Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning

As of 13 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2412.20964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20964 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:12:06.172447Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact1
  • verified fuzzy71
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 580181f6-c219-467d-8a29-3675c4e7e910 · outbound

This paper cites Parallel Vertex Diffusion for Unified Visual Grounding,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Parallel Vertex Diffusion for Unified Visual Grounding,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.723321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.723321Z digest=sha256:bc59639d47763c6ef2e3be5cf25a39042f71b0269cae868d7adbc29fb3b1d98b

Observation e2caa0d2-1939-4c13-b5b1-490a56ac8513 · outbound

This paper cites Align and Prompt: Video-and-Language Pre-training with Entity Prompts,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Align and Prompt: Video-and-Language Pre-training with Entity Prompts,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.729502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.729502Z digest=sha256:8e893d6b81cc557c94c8c742fc04817dbb5bd3d38c3b1b101b04b1ca844d1aef

Observation e24e2e46-65b9-48d4-b2b7-2d2faf469f31 · outbound

This paper cites Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.734526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.734526Z digest=sha256:d242522986d4e740083822fb06c69df9910923b1d4718ef19ced496e6f622d38

Observation 01d8a163-76f5-466d-913b-9e3a768d01db · outbound

This paper cites FreestyleRet: Retrieving Images from Style-Diversified Queries,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning FreestyleRet: Retrieving Images from Style-Diversified Queries,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.740405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.740405Z digest=sha256:bb9497086ad59f792658d5155858b5a526f5c8305b8f4c5b862607a9e33d6424

Observation b1aaa9f2-4362-4779-9054-8aa9998607cd · outbound

This paper cites Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.746306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.746306Z digest=sha256:d50135d9fe057f0cf1f1f9699cf79404e78a4de1336929a7626cef3ad37751bc

Observation 29ef21c2-3c08-41e0-80e4-ea9205a68081 · outbound

This paper cites Dual Encoding for Video Retrieval by Text,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dual Encoding for Video Retrieval by Text,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.751937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.751937Z digest=sha256:49ad52952e4e6b2c17264269743beafac5dd8a95d14f980d39da32ec3acb8579

Observation b8abaa5a-d856-47b7-ad00-344a54ca18ef · outbound

This paper cites Temporal Alignment Networks for Long-term Video,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Temporal Alignment Networks for Long-term Video,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.757710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.757710Z digest=sha256:d65f7705bcfd914e27315d69e2883e231697731b2ff0441e0e71cbb56678d5f2

Observation 9a402677-488d-4e60-b5c4-0642ff5be8a1 · outbound

This paper cites Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.762521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.762521Z digest=sha256:7ac462ba91ebead072a4cc842b4ecb90798402a59ddbc2df07acf6b346fe00b5

Observation fadebf4e-91ec-441e-8b23-2b5dbf0299c1 · outbound

This paper cites An axiomatic approach to the concept of interaction among players in cooperative games,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning An axiomatic approach to the concept of interaction among players in cooperative games,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.767544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.767544Z digest=sha256:a9494131f66478fe0063439c69ec501e4a7435d82cdd527be6fb798ae7dc4a83

Observation dd351e5b-f8c8-4aeb-9b84-439383506c24 · outbound

This paper cites Weighted Banzhaf power and interaction indexes through weighted approximations of games,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Weighted Banzhaf power and interaction indexes through weighted approximations of games,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.772384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.772384Z digest=sha256:9dbee6c781fcb34e83a04a5520624b08909c80ccbbc96e8ce09f698dd144b4cc

Observation f4cecc72-b5a1-4a92-9794-f533cd60633d · outbound

This paper cites Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.776983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.776983Z digest=sha256:1e20fb85e0592d669618d695e21404158c7667de88928a4843be8fd71152b92e

Observation dc667b79-58bb-48a1-b6d1-09c9a51b3d9e · outbound

This paper cites MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.782331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.782331Z digest=sha256:d686b27a9e3c63f87e27e0b5d9922bf52c2f5a807ff2d784c79cdbf944aebd88

Observation 8257d3ac-c9d2-4e89-9ffd-6045984ec663 · outbound

This paper cites Dense- Captioning Events in Videos,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dense- Captioning Events in Videos,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.787122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.787122Z digest=sha256:9ae6bf701e073cd3998c1134b716c874b3801b1c54cf0af7d5d5ce956192b97e

Observation fce22814-90f7-495f-9113-89221bb66181 · outbound

This paper cites Localizing Moments in Video with Natural Lan- guage,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Localizing Moments in Video with Natural Lan- guage,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.791722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.791722Z digest=sha256:025647a6e47f1affb7b025561cf61fe3c58bb977b23efdf9ad1944e51294d789

Observation de3403c7-4884-4a80-9f49-c8750cb0a0bc · outbound

This paper cites Video Question Answering via Gradually Refined Attention over Appearance and Motion,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering via Gradually Refined Attention over Appearance and Motion,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.796656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.796656Z digest=sha256:74c86c67548d344aec503aa76738c77c0436e569049962882b4baf40c3a8e4be

Observation b274597a-7c10-46af-8fe2-9eb2b24caf6b · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.459256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.801152Z digest=sha256:3d92ad80ccc7e7a7529aa4df7be9b62f1b71b6613a0bfe9ea819105cd28eaaf7

Observation 0f155bde-624d-470e-ab62-1c829b685aab · outbound

This paper cites Universal Weight- ing Metric Learning for Cross-Modal Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Universal Weight- ing Metric Learning for Cross-Modal Retrieval,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.442279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.805550Z digest=sha256:e964c19e505b3d5a564977cbb883e140dafa707bc4d9be62b3a535ad9fc52812

Observation 826b2159-8ae8-475c-9961-5f410fb49b6c · outbound

This paper cites Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.426202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.810678Z digest=sha256:7f171e9615c070ecbd5c7b32a761e7681b68b5d6e1fd943deb6d815fd1fe375b

Observation 187ba5fb-4864-40ca-860e-08adcb7f871e · outbound

This paper cites Revisiting the ‘Video’ in Video-Language Understand- ing,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Revisiting the ‘Video’ in Video-Language Understand- ing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.409207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.815304Z digest=sha256:bffa768002b66e05bf7975d716e54e451f5a59e56ba85786e10b29e50148c835

Observation 5131f6e0-6939-480c-90d9-d4f522313cd6 · outbound

This paper cites Fine-Grained Semantically Aligned Vision-Language Pre-Training,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Fine-Grained Semantically Aligned Vision-Language Pre-Training,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.391727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.820300Z digest=sha256:7b22895a97905e75892b690d92e3f4c7b5a67b02f9859b408cafdd032200336b

Observation ed08a169-cf2c-4d57-a05a-1acbe1425b0c · outbound

This paper cites Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.376631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.824893Z digest=sha256:f70bb9b99af347160f14e5027f384d11f4849617709e8dbc121822cb8d042285

Observation be87c5ae-cc08-40cc-b5e0-d85657143a3c · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.360014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.829969Z digest=sha256:74b44d3aa0330ad74f55c3a33e7380e1ea5f0870f4055a8bfa7b39bfe269e47f

Observation ba5899eb-b8d5-44db-9f26-2564ea2e10f4 · outbound

This paper cites Decoupled peak property learning for efficient and interpretable ecd spectra prediction,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Decoupled peak property learning for efficient and interpretable ecd spectra prediction,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.342699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.835414Z digest=sha256:94cd557198e7dd17548b6177fd7226ed95d7299b566b4fcec9f2ad453744145b

Observation 85c7ebb0-03aa-4346-af36-21c2b3c6941c · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.839893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.839893Z digest=sha256:2217a633658841fcf5278d37d03ac015b0be344b12ee090c073c6693afe2f1dd

Observation 0084ff6b-a2bd-4067-9791-3afa3253ddd3 · outbound

This paper cites EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.845207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.845207Z digest=sha256:6dadb987295fe59db67991e09e2f6954687b9e2e911fce14cbbf31e0f9649d90

Observation 7ac6680f-dac9-4815-81fc-e833b5061aed · outbound

This paper cites Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.850062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.850062Z digest=sha256:27ead6831b3c168bd046c61bbcedb5a2580eb434629c5285e0083449d4d5d278

Observation 233a77f0-5372-4513-8faf-0789cd268a65 · outbound

This paper cites Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.326839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.854989Z digest=sha256:65cc53877e1452e872ee86027afafaed80a1a28c8d2f9446c64f49b885da7840

Observation 1457285d-4191-414b-bd55-2cec95ab0033 · outbound

This paper cites Next Patch Prediction for Autoregressive Visual Generation.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Next Patch Prediction for Autoregressive Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.860105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.860105Z digest=sha256:7e82ae17577dcf5f64301c5c7884edad42710767a7d2259ad2723aa7ffe5203a

Observation 06d71f36-46ef-48fe-8df3-c76f22863b9a · outbound

This paper cites Learning the Best Pooling Strategy for Visual Semantic Embedding,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning the Best Pooling Strategy for Visual Semantic Embedding,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.310826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.864794Z digest=sha256:43710641b5a1b5994db370f171cc7da4e040c78e45e2dcd96d5cc5e1bfb85714

Observation 5c620472-2588-440c-b775-3ea6db513a70 · outbound

This paper cites DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.294215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.869660Z digest=sha256:1181c02dfced65793b4b0e7cbd242284ce2207f6aec95d2b66469edfb8bfebab

Observation 60c8fc4d-61b7-47e8-8edf-ed2d186a6c97 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning Transferable Visual Models From Natural Language Supervision,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.278911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.873998Z digest=sha256:f1e4b7b202acb9e24c6c556ceb616f16827a852f099c70b6cb4318d88242069e

Observation 46fe8fa6-37fc-4019-8dcd-858642da83cd · outbound

This paper cites ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.263624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.878311Z digest=sha256:ebf2b832deda350a549cceb21528b9a358e62a03f03e025f5618d4059258ae89

Observation bc757e8e-5304-499b-b4f7-8b5485bbea67 · outbound

This paper cites SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.247617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.883444Z digest=sha256:5d68e325a52413421f70df95de60750d73bc037e0013c69cdae2eab0e5fdccf3

Observation aba342a8-9f5d-4cfe-88f6-5fa7c2d1eac4 · outbound

This paper cites Hierarchical Con- ditional Relation Networks for Video Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Hierarchical Con- ditional Relation Networks for Video Question Answering,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.232752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.887935Z digest=sha256:b2b952515c7641d6408349095efb53e7b6f17b3eec5102f96bc30f09b2c11240

Observation d05302cd-1063-4a5e-9a41-d0d265532e07 · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering: Datasets, Algorithms and Challenges,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.217463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.892548Z digest=sha256:3c8cb573418366d3d4aa6f2ea9afed4eefbb5079aa4135d2fb6fda0abccc5c62

Observation cecd60ef-9e7a-4f28-9cc1-9663220d202e · outbound

This paper cites Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.200092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.897229Z digest=sha256:ae0fb124bd4ea436ae1f77d5cea14f1e571a2bfd5fc63d2d160f666ca753f1e4

Observation 3aee8a22-e517-41c6-8159-c847b1d072cd · outbound

This paper cites Video Question Answering with Iterative Video-Text Co- Tokenization,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering with Iterative Video-Text Co- Tokenization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.183820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.901654Z digest=sha256:659b267db12060767bdaabecdceb74837a8c2b37e6fbf194cc51dbc32079d01f

Observation 3d53e862-8086-4691-8574-2ff90815e13e · outbound

This paper cites Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.167383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.906018Z digest=sha256:d4be14ca62229e5f01b6dde1194ebf6bf9d5e015072f33bf8420fa4a0b5a1282

Observation 6f94017d-14a8-43b1-9f46-e0e0993c33f0 · outbound

This paper cites Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.151121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.910639Z digest=sha256:890dc230673015cbe21c04f8bd3e0e08c9a4c13ba5663a6df31e2ea57b411c64

Observation 06dea930-9905-4701-b330-7647b8ec12b1 · outbound

This paper cites Jointly Localizing and Describing Events for Dense Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Jointly Localizing and Describing Events for Dense Video Captioning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.136079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.915159Z digest=sha256:70e7600408d3a01ac6fd9af65fea8efbf85028b90fd3f22152cc47192684acb7

Observation 72c9e196-afda-4e0b-b89a-ba78e686ab17 · outbound

This paper cites Video Captioning with Transferred Semantic Attributes,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Captioning with Transferred Semantic Attributes,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.121726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.919598Z digest=sha256:d0af374bf4b63320fdca0829524ea6698c897c7f0f7ed0da91d20c09a98181ba

Observation 30e5a4da-b26c-40e9-8b55-4ed14d965516 · outbound

This paper cites Jointly Modeling Embedding and Translation to Bridge Video and Language,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Jointly Modeling Embedding and Translation to Bridge Video and Language,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.106098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.924199Z digest=sha256:f74872be8faff450cc91dc42774ec9a3c8041b45062b1ea06f6d996e0027a9f3

Observation b1c5967c-c4d2-4621-b14f-94ede0299011 · outbound

This paper cites Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.090476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.928690Z digest=sha256:cea613a76ca21e44cb181891b245cff4f839b9a0bef2d24a2f4aaa07ac884781

Observation 7ade93c9-c123-4725-b194-f045060e46e9 · outbound

This paper cites an unresolved cited work.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:12:07.074915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.933482Z digest=sha256:59d15ed7bf80f55d71138a699c888ad520745a1769ae899a3cd9b00c9cc490a1

Observation 07426d4a-e689-4008-9b1a-5ba4e3deb0bb · outbound

This paper cites Optical-model po- tential in finite nuclei from Reid’s hard core interaction,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Optical-model po- tential in finite nuclei from Reid’s hard core interaction,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.059636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.937927Z digest=sha256:65d56335d658b63d851fa056ca1b50ab1839a288fbc56096aec7d1f07cd758f7

Observation 01e9f7a3-a58c-4079-a4e6-1bb56496094c · outbound

This paper cites Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.044075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.942379Z digest=sha256:30b43cb9b031831300244a2d76010dc285f2d34acfa179636f4813805a81e8d9

Observation a32ecb05-ee33-4dca-90a3-2989dc7ab9d9 · outbound

This paper cites VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.026810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.947661Z digest=sha256:af2b7ead4a2bbf22fe8c0b64b64115a135a3bb5585e9192b17a5cdb4c38c2319

Observation d89eefbc-6df6-41e0-9283-65bb44626049 · outbound

This paper cites Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.010019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.952180Z digest=sha256:d38a18a6088ec710acbc0849670a6793161b908d68b83524ba96da0724ce5f29

Observation d12f0efa-572a-49ea-964e-9618a659463d · outbound

This paper cites Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.993715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.956766Z digest=sha256:027b10e8983d9ce9c5c943c646a5db3d213b2447c579f0cac0586b19674256a2

Observation b4435afe-cd1d-43ca-aee0-7ed9b362b0b1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.976856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.961335Z digest=sha256:c86d24b483e702b049e5e1ef5b0ea3cd0f115ea3a264d255ede4cec24c98fda4

Observation 6b280c4e-95ac-4ae6-acd2-1ccd5c61d464 · outbound

This paper cites Kullback, Information Theory and Statistics.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Kullback, Information Theory and Statistics

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.960734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.966063Z digest=sha256:5741b572d065983f9ee1d8fd685071866b3a13c319f15b6cf73e1af99bea675b

Observation e03b7460-3f52-4367-9457-1b99a1dcf37c · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Study on density peaks clustering based on k-nearest neighbors and principal component analysis,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.971077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.971077Z digest=sha256:92280b82e44d8090f2a85c0ab40aa9a01e7d28bd3cf52e06e2124fb0dfa451ae

Observation 5ee34466-8277-4045-8893-f4b3911909ad · outbound

This paper cites ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.934023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.975596Z digest=sha256:6fba409500b38596be7e71ff70077619d1f617dfc04bb97b4be8d083af20e440

Observation e3afa91d-1e33-4046-8e5e-f679205bacfd · outbound

This paper cites Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.917159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.980369Z digest=sha256:2795005e77b3a4f175bc3f5fcbcc7fc066df4e7fd7a4a99fa397aeeca3965931

Observation ba9f6136-6bcf-4314-ba69-d67834dda4f0 · outbound

This paper cites Cross Modal Retrieval with Querybank Normalisation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Cross Modal Retrieval with Querybank Normalisation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.901496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.984936Z digest=sha256:1aa789d8eb6ba19866c9cf8a906d17c80c9be651faf18da4361badcf747c99b5

Observation 997268e8-2bf2-473c-b96c-eae4f7ab6974 · outbound

This paper cites Multi-modal Transformer for Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multi-modal Transformer for Video Retrieval,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.885504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.990474Z digest=sha256:36121d3c1b19b2154ce3193e5686d900b5d1fceb39246b755d274a283151b29f

Observation 8cce65b0-6d35-4cfb-98b6-760e8e60abd7 · outbound

This paper cites T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.870428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.994958Z digest=sha256:8bb496732888cb7a4e1853ee8329481cd8782cb94a2fe30d10682063bcc2f244

Observation 70980822-f817-4cbb-88e1-50106cc2be1d · outbound

This paper cites TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.854606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:05.999483Z digest=sha256:c6e94e502e98948055b8175b6a6be537c33e77134bae72d515ce4078f6d9515e

Observation e55b30e8-aa31-416c-8cce-50fef8ba3127 · outbound

This paper cites Support-set bottlenecks for video- text representation learning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Support-set bottlenecks for video- text representation learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.839480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.004645Z digest=sha256:e359877a811b37ad3699f0b1830e730efe37319c466049a8131f39e5623e0636

Observation bd42f802-5095-4e2c-80d9-bb264a531cfc · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.824809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.009501Z digest=sha256:de795b9ab3c5fe0e18f534325872c0683025f10427bf8140fbd2c7feb2b0c50d

Observation cb85fc88-9c83-4ee4-9f5d-ed059dc9c89e · outbound

This paper cites X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.808890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.013862Z digest=sha256:a8d2b3aab14860ab714fd3e324603029cea6fed9d3db40b0bd1f3e2b3ca9b452

Observation 57368048-0dcc-483a-884b-ccffb54fb35d · outbound

This paper cites TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.793165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.018307Z digest=sha256:e6daa0b5abe3b0a9d60f45310894cf7aae0926c683acd8f2f0f1069e39a56710

Observation 133ea5e1-da19-45f1-a8d2-e39daae0b236 · outbound

This paper cites UATVR: Uncertainty-Adaptive Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning UATVR: Uncertainty-Adaptive Text-Video Retrieval,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.778433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.022809Z digest=sha256:a89b3dabef4ce0c37fa6e87f49755efeefed6db1d2955976e671af50ca4c2cc5

Observation 776c9f9a-afb4-46d5-8000-cf06ec459f02 · outbound

This paper cites Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.762238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.027184Z digest=sha256:cc59aba9ec16ec311973d7dcad0a5117d7b17cdc9a9b8b58ab62b772e6776637

Observation aa1b7248-99e4-48db-94ae-e21fd7fd86ed · outbound

This paper cites CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.746220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.031875Z digest=sha256:c0aa7d44935d5ab7902c8d9708e680e30c2268bd246ee4df18009948c694a5b4

Observation da4aa689-5234-4760-b1f2-2ba1ce0e7b18 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Representation Learning with Contrastive Predictive Coding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:06.036476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:06.036476Z digest=sha256:917759495cb63ea0bfb884c4e6dd46e7ff7a5ef6a8b3bcd6faad4240384913ed

Observation c0fb8e42-135b-425f-b127-0967ae37d3cb · outbound

This paper cites Use What You Have: Video Retrieval Using Representations From Collaborative Experts,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Use What You Have: Video Retrieval Using Representations From Collaborative Experts,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.731010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.041258Z digest=sha256:9f8a2ddea3576cd9cf49350d5f6a69fded033c69711e15001af42003f7e38f8b

Observation 79c510ef-24e5-44a9-a331-f0b7d7605a6c · outbound

This paper cites HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.715584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.045870Z digest=sha256:c93d85d4508108c8df25a27c7367c872e7adacb5941bd0e525467a330e7f676a

Observation 427b624c-1c82-47b8-8a34-34b1837428ce · outbound

This paper cites A Joint Sequence Fusion Model for Video Question Answering and Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning A Joint Sequence Fusion Model for Video Question Answering and Retrieval,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.700240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.050286Z digest=sha256:159f90809cf91a7b76471e36d440a089ebf57ce305c4279295d5dfa6170b560f

Observation b00737e7-0ff0-4a41-a683-c3b4ddc93247 · outbound

This paper cites BLEU: a Method for Automatic Evaluation of Machine Translation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning BLEU: a Method for Automatic Evaluation of Machine Translation,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.684401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.055538Z digest=sha256:f619ccdafd91d8b80ac990b50d2ff92ce913a16a57adcb8fc50b5bb6962b7b17

Observation f6780819-c2ea-4a2c-ac1b-75b072106f9e · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:06.060027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:06.060027Z digest=sha256:e0ffb71b356e549e4c60ca6eafadfad4d075ba59d24964d300a7b4b88ca59a00

Observation 2608b2a6-8f9d-4e4e-aa70-44dc6abe8f24 · outbound

This paper cites METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.669554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.065222Z digest=sha256:36be436ff084bfb61ecc9df886fd9072ed383df3e22f9883e25ead1db4766a80

Observation cc57212a-8e67-4c84-9fdd-f2fe437e7b9f · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CIDEr: Consensus-based Image Description Evaluation,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.653153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.069878Z digest=sha256:4ee9ec6ad5d7356cf99e1bfc0494b5174388ce6eb14d654e059b71a51433dafa

Observation c880074f-2595-4c3a-8a4b-a8fa0607d1bc · outbound

This paper cites Adam: A Method for Stochastic Opti- mization,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Adam: A Method for Stochastic Opti- mization,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.637445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.074370Z digest=sha256:536c01294d377124ca73963db1c648a8da089da18f351d3c28bd3009ee04d56a

Observation d62a870f-b1c9-4696-8aba-208a5cf6ce28 · outbound

This paper cites Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.622128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.079203Z digest=sha256:d8850c79bdb2301d83912df69b9ade5ddef7687bfc3bd4bd2a62ed689d50f482

Observation d72011aa-5ba8-428b-93d9-d17126a0a052 · outbound

This paper cites Approximating power indices: theoretical and empirical analysis,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Approximating power indices: theoretical and empirical analysis,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.607028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.084651Z digest=sha256:50ad44a1d706ab200e32d15fb890ea9a315b8f938c337871e80ce508fd05ebcf

Observation 8a5e0918-8873-4beb-9cb1-3a340a0a809d · outbound

This paper cites All in One: Exploring Unified Video-Language Pre-training,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning All in One: Exploring Unified Video-Language Pre-training,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.592171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.089292Z digest=sha256:9a26a1940f77d517484c058e22d65522ca084987ac1d563c303060df4341a0ef

Observation a4b19c3e-4df0-4fe8-a0b4-93250bc684a1 · outbound

This paper cites Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.577862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.093859Z digest=sha256:7497203b7e740ac156c92493f37bbfcaf490fbdc124cccd08628fab479d90d24

Observation 59788f69-d0ae-421c-b8f0-aad0845698ad · outbound

This paper cites Multi-Granularity Interaction and Integration Network for Video Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multi-Granularity Interaction and Integration Network for Video Question Answering,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.562922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.098477Z digest=sha256:ce255fb9fb9b9dd7cf6016e446d57f73e9d57e9cbcc542716fda2c8f53856c64

Observation 9dceb0e7-669a-4a6c-9a09-f5cdbcb84564 · outbound

This paper cites Invariant Grounding for Video Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Invariant Grounding for Video Question Answering,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.547762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.103019Z digest=sha256:891eb6eca534f8ef77bb5aa309daf79e103920d45bc8babc262c473ea909a8ca

Observation 03e73832-5573-4e38-a027-897810212dc9 · outbound

This paper cites Learning to Answer Visual Questions from Web Videos.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning to Answer Visual Questions from Web Videos

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:12:06.218997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.108321Z digest=sha256:9bd3611664993bb33ac85d999c6404968369b0a7c0083dc00bbc7d9dcc5fc3a7

Observation 6bbc903e-83ad-4f64-a7e0-0a5b23d8fcc3 · outbound

This paper cites Video Question Answering With Semantic Disentanglement and Reasoning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering With Semantic Disentanglement and Reasoning,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.532572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.113224Z digest=sha256:e77aacf9ce853c872847987b584142cacd678093efab702e552546e96d9a64e3

Observation 87fc9f50-ec22-48ff-8463-251123d092c2 · outbound

This paper cites SViTT: Temporal Learning of Sparse Video-Text Transformers,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SViTT: Temporal Learning of Sparse Video-Text Transformers,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.516916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.117745Z digest=sha256:3b35493b7dcead7b518db687c0115a7675e00cf218b706db5fda2fb91d5164e5

Observation b1503c78-c126-43f8-8679-bb33c5a59b99 · outbound

This paper cites TG-VQA: Ternary Game of Video Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TG-VQA: Ternary Game of Video Question Answering,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.501948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.122401Z digest=sha256:edd770402367a9c3d6d7ec6052f8244e1f4be223381c736dd9ae224ae9ba3c52

Observation eef74904-de05-42e9-a437-f6a71b29b991 · outbound

This paper cites SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.486822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.127080Z digest=sha256:2f97699ac00911fa5e3350abcffe298a6b13f5aca295001ee46c5dfe30dcabec

Observation 0a4ac961-06e3-4e5a-b2be-fe544deeec82 · outbound

This paper cites End-to-end Generative Pretraining for Multimodal Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning End-to-end Generative Pretraining for Multimodal Video Captioning,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.471147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.131507Z digest=sha256:d3b6d0d1083cc33371652b451c5649eba53c515ab816fe1bd2606ec4d4d644e9

Observation e5f9d9c6-66ce-4ac3-b6ef-64bcf0c1f83a · outbound

This paper cites Motion Guided Region Message Passing for Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Motion Guided Region Message Passing for Video Captioning,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.454914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.136069Z digest=sha256:7783354624a6504e3b54d1422a19d82f19fd74b102ee6bc09908307d10eeca87

Observation 49ad36ec-6710-4eed-85f8-b9f6df6b9bd7 · outbound

This paper cites Open-book Video Captioning with Retrieve-Copy-Generate Net- work,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Open-book Video Captioning with Retrieve-Copy-Generate Net- work,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.439692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.140645Z digest=sha256:eec52fa8ab81c454bbca08191988cfb224eba99b242093ee82dd648751733af4

Observation 7cffd5d3-dc94-43ff-802b-23f791be790c · outbound

This paper cites Attentive Visual Semantic Specialized Network for Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Attentive Visual Semantic Specialized Network for Video Captioning,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.424464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.145153Z digest=sha256:69b7140f299ed910b4831074f6a52cf5a5687c919b5a445ccffa671363cb7e29

Observation 6bcd28a6-c000-49e9-98c4-a62a84bc9b9c · outbound

This paper cites Global semantic enhancement network for video captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Global semantic enhancement network for video captioning,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.408778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.149625Z digest=sha256:c723ee73a2a39d1f56958c98d4be65ab994452063cb1fe0d7723899ac755f5b1

Observation 5b245349-1528-467a-a873-cb780712c22e · outbound

This paper cites Accurate and Fast Compressed Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Accurate and Fast Compressed Video Captioning,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.391481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.153955Z digest=sha256:9bd962f997c6635de30db713d4417deb9ad6ce1354717d925161cab8e12a1b55

Observation 5356a222-97a4-4788-8cd2-56edf440d795 · outbound

This paper cites Emotional Video Captioning with Vision-based Emotion Interpretation Network,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Emotional Video Captioning with Vision-based Emotion Interpretation Network,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.375787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.158851Z digest=sha256:7f29bc6375f78ddbb26b103ed094c3b33a5fa14a958c840f0a324b95a5904f2e

Observation 989e1414-4faa-433e-81f7-9d43588b5c27 · outbound

This paper cites Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.360802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.163381Z digest=sha256:1321264a7b99e905bfc4cd6e9e7d78ed24ef3b4738b860552c0ddf7872d2af14

Observation 2b4da795-3272-4e61-94e2-545790861b51 · outbound

This paper cites CLIP4Caption: CLIP for Video Caption,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CLIP4Caption: CLIP for Video Caption,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.345385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.167875Z digest=sha256:c7e7c7fdf3a1cbe28d3b69baee3356835c2127f767017b1e5d3882e0e63364d8

Observation 44183826-f439-4dc4-baa9-53623641cdf6 · outbound

This paper cites Visualizing data using t-SNE,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Visualizing data using t-SNE,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.330034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:12:06.172447Z digest=sha256:b20cca01258371bd1986b62e6c09d1937d26b14c23a11e6edab989448b73efb6

Pith citing papers

No inbound Pith citation observations are available.