Pith. sign in

Paper Citation Record · LEDGER

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning

As of 12 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2412.07704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07704 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:40:22.275159Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact0
  • verified fuzzy55
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 675d14c4-d3a2-45a6-a40d-716825cb5614 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.730017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.730017Z digest=sha256:5f3857c0cd5bad5190d55fb90b26e3e44ef7959bf7224295627be029bf053a20

Observation b47675f9-c508-4379-bb3e-d4d290741054 · outbound

This paper cites Vivit: A video vision transformer.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Vivit: A video vision transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.736471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.736471Z digest=sha256:8d0bcfc4ce2d12b83ecab9c07e21bbf9362e6c329f947a3c3b28d7ff74a99226

Observation 38d5c2e2-3795-476a-b98c-d71670add427 · outbound

This paper cites Hiervl: Learning hierarchical video- language embeddings.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Hiervl: Learning hierarchical video- language embeddings

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.241716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.744196Z digest=sha256:5c5a53d15dbb308198936a5906f5cce27db26d1f7be6ddd0ccb0e8105b0ead49

Observation b74dc9e9-fa4a-41c3-9209-ef3c5f19b86d · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.751687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.751687Z digest=sha256:a8e1fae1dc14949eb8e7ae9cf4b998afb329f92101b309b98821a25889b5613a

Observation f5fc93dc-7da2-4886-927a-590bf0e0459a · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.215387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.759600Z digest=sha256:99b1344c5b875dd6e39ee5aeeeebac33267dd845079b245e747501e449741850

Observation 34c8f8ad-7228-4930-8442-ff00864656f5 · outbound

This paper cites Blinkdl/rwkv-lm.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Blinkdl/rwkv-lm

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.192917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.767273Z digest=sha256:d19363c5b556421d6a65dbb33674ff47038acc32704d9a9161dd0d2e0c768f01

Observation 172b787c-158d-4dea-9fd9-a028f716f0fc · outbound

This paper cites Revisiting the” video” in video-language understanding.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Revisiting the” video” in video-language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.166350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.778008Z digest=sha256:d5192281a15767b1ce5bcb4e1420f4fc614366d0566406b3e0bdda25cd0dc686

Observation 4bcf2b93-2c01-4852-ae96-741a45dfa03e · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Activitynet: A large-scale video benchmark for human activity understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.787702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.787702Z digest=sha256:f3c19731e372ce27c476b672b17f1b820c11b9a082ed902a75ec11637c3ce16f

Observation f0f99526-3ada-45a8-923f-01379c9e66b3 · outbound

This paper cites Locvtp: Video-text pre-training for temporal localization.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Locvtp: Video-text pre-training for temporal localization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.132133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.795890Z digest=sha256:d7637f164856909736809d48f9888c6ad836619f7bd746828574313b723f595a

Observation 51d10e7b-49f6-44d8-ac88-ea40bf174a17 · outbound

This paper cites Metaxas, and Hongxia Yang.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Metaxas, and Hongxia Yang

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.106905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.804376Z digest=sha256:d0d486eb396a4f69207de7716af421a9f912e53b997bb498c4421fc2aeb31fb3

Observation 936c7b90-1738-482d-b743-2c656a516604 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Gonzalez, Ion Stoica, and Eric P

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.086270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.809265Z digest=sha256:2afe2c3e081f2ab9fc47cab99a692214343d03f9294e34f9f03afe2070d98502

Observation 83564cf4-7ebe-42db-ba93-4a78643a4965 · outbound

This paper cites Unsupervised and semi-supervised domain adaptation for action recognition from drones.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Unsupervised and semi-supervised domain adaptation for action recognition from drones

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.059920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.815871Z digest=sha256:b881c49ca59c6c06486550dcc7238062cd1867fa2526dfdeb4ae055e41e64278

Observation e0c386f9-b55b-4cae-92fc-1303633d524b · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Free dolly: Introducing the world’s first truly open instruction-tuned llm

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.039299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.822659Z digest=sha256:97fe424e0024bdf9620880b5705bad32d949e8a629f32021b7663b34ea30002e

Observation 18e3b570-4d85-432f-bdcf-5c61fa1298c8 · outbound

This paper cites Prompt switch: Efficient clip adaptation for text-video re- trieval.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Prompt switch: Efficient clip adaptation for text-video re- trieval

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:24.011454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.831924Z digest=sha256:28f2f5e015a42cc5b4cd59fbd3c17bdf0be2929b2063f275e0dc273c94fcebc3

Observation ea1e73ca-d9b1-40f0-abc4-c90b0ae38730 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Imagenet: A large-scale hierarchical image database

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.837845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.837845Z digest=sha256:0e6e8e1c1f12e1a8d02b45d92b11007bdfdeeef234e672134f36c7ba2a78186c

Observation 44711116-d431-413a-b5ab-30eafca01d55 · outbound

This paper cites Text-guided video masked autoencoder.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Text-guided video masked autoencoder

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.967912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.844329Z digest=sha256:5982defc95b6f85b90d098799866b02b2cafd1cfb755f1a5c1cb27cc7c473905

Observation 87dcb7e5-5f7f-4c77-8885-fc4d1ed888e5 · outbound

This paper cites Slowfast networks for video recognition.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Slowfast networks for video recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.857535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.857535Z digest=sha256:e69e44fc90d72e3bf8781775c3ca0e8f095efb836fc2eb75a898822bfcf8961d

Observation a19e7fb1-7fe3-4e50-992a-7320ce51d2b9 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.865905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.865905Z digest=sha256:2f5122b1825703567d5b49503f2052bdeed767d019c833bdbc607731af3d5a1c

Observation 18b4b722-8935-494a-9cf8-8deb5bf16990 · outbound

This paper cites Multi-modal transformer for video retrieval.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Multi-modal transformer for video retrieval

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.922208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.873145Z digest=sha256:cc2d821044088aca2dd34a5cca15df4425d0ee0de6670714f60ad3722d86fca7

Observation 080c0049-af7c-4442-88b9-29b1076e3965 · outbound

This paper cites Openllama: An open repro- duction of llama.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Openllama: An open repro- duction of llama

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.881286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.878555Z digest=sha256:1805438ad6a95571db3d1b142122b2ba4005189d0cfc299692d7a7b762ab66e8

Observation 858f3263-772b-419f-bf27-ab7e57e7adc8 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.887912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.887912Z digest=sha256:5df427b1c4530c05a04e275f491c7e2cd52e00db6a729db3201769122ba2b242

Observation 1dc6f2b7-473c-420b-83a7-58515e2278ce · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.839314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.893595Z digest=sha256:7d0710607027f6beb13f1c9af8cfb4bc0403f5edd8c75eb8980f24ce2a39a72d

Observation 0ec6fe47-169b-45ab-a838-17919f451dad · outbound

This paper cites Long movie clip classification with state-space video models.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Long movie clip classification with state-space video models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.818116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.901974Z digest=sha256:f61f8dd4e98bb440e9bfda2d05861d61bb576b987538a4e2497c2f636d2f51bc

Observation df4f55af-fa3a-40f4-8e0a-65cc1978df08 · outbound

This paper cites Video re- cap: Recursive captioning of hour-long videos.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Video re- cap: Recursive captioning of hour-long videos

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.792138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.908126Z digest=sha256:8bfa696bde0a965fff40ed0510db013a64ae40549d4120265fa3681f6ab89d90

Observation 363a9f56-e2ba-483b-ae32-7a3e6db037f0 · outbound

This paper cites Perceiver: General perception with iterative attention.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Perceiver: General perception with iterative attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.914293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.914293Z digest=sha256:a2f654fb066bd764d080f06d942609a8e48beab98324acd589837ec6f24b5c0a

Observation 812c4982-dced-47e4-85a8-6fd4b60f6468 · outbound

This paper cites Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.761987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.919118Z digest=sha256:1505be37b5bc3f73d5a79b304642efb7af22c1ff44e1989d410461f12a1ecf19

Observation c8dccd7a-5203-4e85-a82d-112f67b55425 · outbound

This paper cites Dense-captioning events in videos.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Dense-captioning events in videos

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.743632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.929454Z digest=sha256:6a782ab3b0b51e3650784591fc993aafb5e075dd58c6f76100073db9144b4596

Observation 986c2357-dba8-4c17-8413-b75dd961f300 · outbound

This paper cites Video Token Merging for Long-form Video Understanding.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Video Token Merging for Long-form Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.935406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.935406Z digest=sha256:083fc32481a07edc7327a4ca07235ad91ad867ada1a7ab2c838fe4783d006f21

Observation 55c1a9b1-a921-4da6-a56e-5fc612464cf6 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.710374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.940599Z digest=sha256:23120e63b0578af77150b5506f026ba5b1548848e1f871f71c45e892aee5753b

Observation 3fdf41e7-e82d-41c5-a7b3-2187ec78968e · outbound

This paper cites Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.692431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.949222Z digest=sha256:a8e41807ba00f04a33c5c6d78be30a5516233e5063b8ea5705120ee6d4811612

Observation 1a756374-a7bb-4303-8d5e-f8fd96dfb66f · outbound

This paper cites Hero: Hierarchical encoder for video+ lan- guage omni-representation pre-training.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Hero: Hierarchical encoder for video+ lan- guage omni-representation pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.673939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.954005Z digest=sha256:2d1670ab060ceedbf3097801d0917cbaf1524757df29747eefe97f49bc70e9db

Observation 22634687-060e-427a-bff1-b6c0f8143e31 · outbound

This paper cites Lavender: Unifying video- language understanding as masked language modeling.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Lavender: Unifying video- language understanding as masked language modeling

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.648567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.962406Z digest=sha256:5f8af59f3cd98d9bc669b2ace5904b8ccae192b01b968afb572c0048ac441db7

Observation f49ba053-96f9-450c-9b42-0083c9bdfcd9 · outbound

This paper cites Ego-exo: Transferring visual representations from third-person to first-person videos.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Ego-exo: Transferring visual representations from third-person to first-person videos

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.615730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.969173Z digest=sha256:5f02e8d6ec77973e7256a4d4892e70d05e52c8bdad8d408d7afee7ce3211297e

Observation 0de33a2d-d776-4a50-a48b-0f03f32bb807 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Rouge: A package for automatic evaluation of summaries

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.975946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.975946Z digest=sha256:c60673d3e07d31a2721751c9a8b346eda0732438d3667de6d4929ebc98329ea7

Observation e17bd6b1-8794-4626-8b68-7eb329baeab3 · outbound

This paper cites Egocentric video-language pretraining.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Egocentric video-language pretraining

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.564693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:21.984672Z digest=sha256:a96a19a739187c8396f2c56176997032dc78ef090d23dba220fb9d43bd46009f

Observation 8c889a7f-4fe5-4376-b0fc-a079288a3226 · outbound

This paper cites Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.536749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.009102Z digest=sha256:4caa7b85c757f3afd75226115ccfc992651e61857ba6a2f9760ad177a2e82e53

Observation ee9993b1-19a1-49d4-89a9-784672706038 · outbound

This paper cites Sgdr: Stochastic gradient descent with warm restarts.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Sgdr: Stochastic gradient descent with warm restarts

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.523162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.013950Z digest=sha256:9d68f56b3af73f3eac19f21a4210772dd34bc040c85dda507c3cf0bd82b3bc63

Observation 78a68820-7c24-4602-8e0c-86c4c0a60891 · outbound

This paper cites Decoupled weight de- cay regularization.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Decoupled weight de- cay regularization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.492004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.022841Z digest=sha256:4ae67e78e3181b91d40f23a8ec6b2e5025d847006455f869dfca353b81daf378

Observation a23ca64a-d0e1-4f27-9215-6301de439714 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.461142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.029958Z digest=sha256:6af8c4c28391c1aa1052986f5decb7e7a2c607f98722c9b3f12719d53b4247ca

Observation ee2f2e22-5a24-44d1-a06c-883deee273bf · outbound

This paper cites X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.441679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.039956Z digest=sha256:e7019658b4aeca51ac33bf7a08db02b14cc5029bf73093260c91172d99f9b606

Observation e2abef21-70a8-4a9c-bfad-dc77572e9c03 · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.412625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.045764Z digest=sha256:6446da84fed00e8f152d7bc897674da1108e1884405adbebab31ae049720c5b8

Observation 8e2154c4-c4e3-4eb8-8a58-997ada5779eb · outbound

This paper cites HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.374584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.052408Z digest=sha256:30395cfec556f07cd6e1d7b87829b61db7a4f72d2b941dafbf627571aae4142d

Observation dc123eb7-8a58-492a-a5d1-cb67a5a964a5 · outbound

This paper cites an unresolved cited work.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:40:23.343505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.059032Z digest=sha256:db081be6f12d83a98ac105c419e8a1b13da26f4f16e738f8555ae32dd6499ae6

Observation cef34cec-a4e9-4292-85e6-6e4926a25ecf · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Slip: Self-supervision meets language-image pre- training

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.320171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.064479Z digest=sha256:7518359475b0e17fe9ee30b460dee41172c7c7b53291f650512da0edebd50b47

Observation 79a35dd4-7a64-46fe-a56b-0bc4207a3e97 · outbound

This paper cites Keeping your eye on the ball: Tra- jectory attention in video transformers.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Keeping your eye on the ball: Tra- jectory attention in video transformers

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.297289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.068892Z digest=sha256:99e47a4e5f1cdd49103dcba9f61fc47bf5420065803dc2ea92b8065037dece0f

Observation 201ee4fa-1134-4e8b-86f7-fa24e8dc2366 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Learning transferable visual models from natural language supervi- sion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.121481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.074970Z digest=sha256:3e61bd3474c8ee62dc1e5436ebc5a1d6f4fb0e59ed5c1b78308e1ee503c06172

Observation 5caf0299-3444-4b8c-a6ab-42d316724101 · outbound

This paper cites Movie description.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Movie description

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.098320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.084101Z digest=sha256:ecc42c6e74ffc9234ca8674bc1793a271aff2125a5f147fe160dc63c0dae1721

Observation 97dabd7e-7888-4e23-873a-3567774d950d · outbound

This paper cites Actor and observer: Joint modeling of first and third-person videos.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Actor and observer: Joint modeling of first and third-person videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.076278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.092041Z digest=sha256:b225fe673c93947ffd41fc473f933b3766d606bef5cf77e2674ee3d834705007

Observation 1eccf011-45ad-4296-a59b-3a6c2129d3e3 · outbound

This paper cites Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.101057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.101057Z digest=sha256:fae92ff0a6d8e5b0731fb4b10908a9a04406df55aa677558da9a3094de8927a1

Observation 5bcd8913-db3f-42d7-bdf0-bd88170d5c2e · outbound

This paper cites Videobert: A joint model for video and language representation learning.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Videobert: A joint model for video and language representation learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.057996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.106058Z digest=sha256:e62e12f89987a57bb0474fe1487aaff78a5d6725bbf1ca5e3753ce8e1d8b377b

Observation 035e5820-11a0-45af-b166-02439be83035 · outbound

This paper cites Long-form video-language pre- training with multimodal temporal contrastive learning.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Long-form video-language pre- training with multimodal temporal contrastive learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.038963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.110282Z digest=sha256:f2a4317396ff3774f745541915e59d2d384d1845af52a98696069a0bb2530eb6

Observation 4b7db35b-33bb-4347-990f-56ed75b03109 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:23.018025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.114774Z digest=sha256:aa09625a033d9712bdd66e6def834ceb95d7e117aeaac0b9da7b2906ac57a873

Observation 5354a18e-c39b-4408-b1c7-d3eba3df0e71 · outbound

This paper cites Perceiver-vl: Efficient vision-and-language modeling with iterative latent attention.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Perceiver-vl: Efficient vision-and-language modeling with iterative latent attention

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.994481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.119567Z digest=sha256:8ed43c3af2b224b97311ab6f60db3c02d3246032cb1b1257888bc23388b0a92f

Observation 12a33b53-98b0-43d8-b5df-db60e4f26e4f · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Internlm: A multilingual language model with progressively enhanced capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.127678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.127678Z digest=sha256:3a709df87f046d8699f37063a8e93dc369c38dfbcb7d8cab43d76821ab55cdf5

Observation d1d96876-a14c-44df-889e-60819aed4134 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Yfcc100m: The new data in multimedia research

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.960559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.132114Z digest=sha256:51ce084a094494c2aaac78466605561d430f704a924bccb27166bdcf079fcab7

Observation 2bcaf4c2-a33b-4e66-84dc-5b170c1722b7 · outbound

This paper cites Attention is all you need.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Attention is all you need

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.136523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.136523Z digest=sha256:ffb19a4c2ffa5b5a07e3d5025cf0dbf911bc55c390d0de4e92296a34d237c087

Observation ab75073f-9ea2-4f39-bbb0-234ffc962ceb · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Selective structured state-spaces for long-form video understanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.928282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.145845Z digest=sha256:9f5512fa883345e06d05421730d5fdb73448708374fdba492b2c24e393638e43

Observation 4a3faca8-1c04-4fe0-827a-19d99d598132 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.150840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.150840Z digest=sha256:3e92501f12c73c26dacfd0b6e291eb7d1d778a348a7185c4365ae39706b89b22

Observation 73d944d8-079b-4d01-a357-36b6102fedcc · outbound

This paper cites DHP Benchmark: Are LLMs Good NLG Evaluators?.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning DHP Benchmark: Are LLMs Good NLG Evaluators?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.157060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.157060Z digest=sha256:1fb9e0d17a2135552191842aa97db04f6d694aa2e5c467b64d3905486c3be67c

Observation ce0bbc7b-c8f3-4d86-9372-7f6bd640d9bc · outbound

This paper cites Unified coarse-to-fine alignment for video-text retrieval.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Unified coarse-to-fine alignment for video-text retrieval

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.907410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.169841Z digest=sha256:c4dae30cf844c6468a07b768b327f5485efd358506e7bddf0391ce9e010735cd

Observation 0217e50e-9914-4bfe-a08f-03f1b55941d3 · outbound

This paper cites Towards long-form video understanding.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Towards long-form video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.887915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.179293Z digest=sha256:ef5fad3a2e81626dee12cf71898bc4a4948be90646d07b3b6c41a32806113349

Observation af4a49fd-e24f-4298-a8e9-0d1986020330 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.188066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.188066Z digest=sha256:bea062ed7cf2025e3ad56efebdb1daf2899ad7b4a32a1f595f40112be7ec590c

Observation 8634cbef-9b21-4947-afe3-a6ee2c5d0267 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Msr-vtt: A large video description dataset for bridging video and language

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.853043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.195527Z digest=sha256:c3d0a7193ac686a1437966173410e7fa4f89433058251da0239f02ef5556d369

Observation 12f45cdf-17c8-487c-b5ef-6ee4bd4c4d86 · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.808404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.199428Z digest=sha256:eca66384b3ea448bac302151764be5b14fa436457bc3b2b3104771d08167c0ae

Observation c95dff12-9dd3-4d14-a1da-b3f0b253e691 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language alignment.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Clip-vip: Adapting pre- trained image-text model to video-language alignment

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.788021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.203386Z digest=sha256:73164949fc0a31b2506df7710718ee23af441ee6573ca41135235785dc575105

Observation fe072311-6922-424c-8289-b740a55c263e · outbound

This paper cites Just ask: Learning to answer ques- tions from millions of narrated videos.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Just ask: Learning to answer ques- tions from millions of narrated videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.212547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.212547Z digest=sha256:1bca4de8a294615b91080481687d2a1040c06aaa8c2d826525f1984a7d0d9680

Observation f430c376-b266-4521-b89d-a5dc611e5127 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Zero-shot video question answering via frozen bidirectional language models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.218117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.218117Z digest=sha256:3f9a6624f4df0bd181726c993d8f9bbc712f1a74f4c6534ddf92ffb58c5172e4

Observation bb140d43-9d0b-4397-9990-cf86e245ea77 · outbound

This paper cites Taco: Token-aware cascade contrastive learning for video-text alignment.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Taco: Token-aware cascade contrastive learning for video-text alignment

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.735572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.223480Z digest=sha256:224d9fe03ee1283ea9675aa17c9f161a856b2d6e255c8bacda38a30195867935

Observation 37292d9b-0b3b-4440-8f52-346e291a77df · outbound

This paper cites Scaling White-Box Transformers for Vision.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Scaling White-Box Transformers for Vision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.228956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.228956Z digest=sha256:aeeaf51af2be790b992dfda527c9bb670d3f0a3b8a7b0af1aa47cd613cfbd004

Observation b577bacb-e363-45ca-a5be-e2f1cfe3a8c4 · outbound

This paper cites Filip: Fine-grained interactive language-image pre-training.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Filip: Fine-grained interactive language-image pre-training

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.713450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.235771Z digest=sha256:2a6962a0600b21acdc0ad438fd923a462637ed6353b65270a498363f87ad9578

Observation a2ed1c0f-67f9-43c4-815c-4c1447826737 · outbound

This paper cites Hitea: Hierarchical temporal- aware video-language pre-training.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Hitea: Hierarchical temporal- aware video-language pre-training

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.685272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.241823Z digest=sha256:28f14146908d52ced69e04338668bcce31f7ba076b1cb5f609db1498f7ca568d

Observation 5dcc4eb2-2123-4225-b95d-e4cab5cdf4cb · outbound

This paper cites Self-chained image-language model for video localization and question answering.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Self-chained image-language model for video localization and question answering

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.246253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.246253Z digest=sha256:8ca81f822449224baa49055ec0894815a8d334c5b671433d5d5a2737d5ee0e7a

Observation 45c29bbd-b6da-4944-a25c-ea7bac2ae2ef · outbound

This paper cites Learning from inside: Self- driven siamese sampling and reasoning for video question answering.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Learning from inside: Self- driven siamese sampling and reasoning for video question answering

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.654086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.251821Z digest=sha256:d8d8c3363c4b4518c12001093de56392b55688b698084229e05c8a3a5f3f09eb

Observation 78cc4f21-adba-4b55-b2e0-e883298bb3e7 · outbound

This paper cites White-box transformers via sparse rate reduction.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning White-box transformers via sparse rate reduction

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.634547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.256512Z digest=sha256:a478a6a4ea8970d2868f396589f4653c53c980f634bb19d8eca5fd1481deaf90

Observation f07426f0-cd41-48c7-9ff9-d8d981f15060 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning BERTScore: Evaluating Text Generation with BERT

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.261203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.261203Z digest=sha256:6d3776de1ba243252f080c07368085a7f07fe0137b2f3d07f6123d6b4a0b8cd6

Observation 1946b448-e8db-4ff6-9c31-a3e64bfe55ab · outbound

This paper cites Videoprism: A foundational visual encoder for video understanding.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Videoprism: A foundational visual encoder for video understanding

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.608376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.265966Z digest=sha256:c4574cb73560fd49888d3a2d3629d2d34dbca86de8972972abaf09a4bf625204

Observation c9a9235e-43b2-49bc-a573-5842e84d09f8 · outbound

This paper cites Cen- terclip: Token clustering for efficient text-video retrieval.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Cen- terclip: Token clustering for efficient text-video retrieval

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.591220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.270685Z digest=sha256:dd283430030ce89159ab35084737f4a75c43442a66b9c261c13a38c3da6b0d2c

Observation d237a457-aeb9-4f15-a5ed-e7d2a9bc79ae · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning P Xing, Hao Zhang, Joseph E

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:40:22.568705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:40:22.275159Z digest=sha256:b0a772f7c7bddb8f8872e030f3f831c110cfa6c2d155d527ffdb359de1f5f8aa

Pith citing papers

No inbound Pith citation observations are available.