Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

As of 21 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2508.04546.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04546 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:00:03.806881Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8c106e4-32ed-4feb-a878-f068044377f2 · outbound

This paper cites A new surveil- lance and security alert system based on real-time motion de- tection.Journal of Smart Systems Research, 4:31–47, 2023.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding A new surveil- lance and security alert system based on real-time motion de- tection.Journal of Smart Systems Research, 4:31–47, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.376583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.626468Z digest=sha256:a8682ef8e91c32aae88e4a5b584ec25df1a18cdf5f1e063bed922fa80a5640af

Observation 37c3ec9e-6e40-4b02-8e1f-346f489af6e7 · outbound

This paper cites Miniroad: Minimal rnn framework for online action detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Miniroad: Minimal rnn framework for online action detection

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.368076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.630495Z digest=sha256:2586583ab58f0230f26a6ce1a8f94ec9a40a055b02acb90382bd4ef25a1d2394

Observation 4102aee0-96d1-4b0d-afb5-ce00ecd9c7a2 · outbound

This paper cites E2e-load: end-to-end long-form online action de- tection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding E2e-load: end-to-end long-form online action de- tection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.359095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.633772Z digest=sha256:1820a97748c96799c47ce4a7abbf316af9e034503aeee78238c208de1390a359

Observation 68de790d-8497-4e7d-998c-dcd1201b4843 · outbound

This paper cites Gatehub: Gated history unit with background sup- pression for online action detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Gatehub: Gated history unit with background sup- pression for online action detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.350553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.638144Z digest=sha256:c5f22a470d1dbc8fe3e9b8c4d0bcebfca7153732ab62fa2e4906f406f79a24a8

Observation c63d8f15-d39a-484b-a3b9-b56a872f5d50 · outbound

This paper cites Enhancing Long Video Understanding via Hierarchical Event-Based Memory.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Enhancing Long Video Understanding via Hierarchical Event-Based Memory

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.642025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.642025Z digest=sha256:6afcc0efbad75ec2783ec18402999abe9e870ebbf90d584fd6e1d8b3ffac6ff4

Observation 5766dccb-66ec-4831-9344-7f1a5f42bcf4 · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.645941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.645941Z digest=sha256:21aeca6a2d36854f9645f070644bd604d59677d1ffc461adf1fe728317f9ef37

Observation f4b83441-b569-49c3-9940-ef9418adec9c · outbound

This paper cites Moment detection in long tutorial videos.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Moment detection in long tutorial videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.341043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.650510Z digest=sha256:721cb4c96628d96dae25e47ae21858ec5027512506abb26c0e098576218772f2

Observation 4d64475c-9df3-4c21-8876-73943611cce8 · outbound

This paper cites Online ac- tion detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Online ac- tion detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.330770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.653383Z digest=sha256:e8e4d661abcfd5b4222bd177e834cf9e0bf86367b3b6202b0da6d389949f1cd0

Observation e3a56d9d-92f6-423a-82f3-9463e0ace12b · outbound

This paper cites Learning to discriminate information for online action detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Learning to discriminate information for online action detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.322006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.656452Z digest=sha256:7f95e9b53c37a65ee510250aad9a8f2cec801e1c14c5a9042aebedfa6dd2379c

Observation 9ae8c4fb-bdc2-4f11-bae0-81e5fe594f52 · outbound

This paper cites Temporal sentence grounding in streaming videos.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Temporal sentence grounding in streaming videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.313071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.659145Z digest=sha256:8d9f2072f5726048c037635ddb8340be840a971ed72cdffbfa1cfe14a41ba8cf

Observation 1bc21709-7208-4148-878d-b399b0cc003b · outbound

This paper cites Tall: Temporal activity localization via language query.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Tall: Temporal activity localization via language query

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.303986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.661894Z digest=sha256:0001996a3ee6088708a8d224d61e026b8f5f7a46033cd99600e9d362ce3dfaa5

Observation ddbaed9e-9cbd-46e4-b495-5ca378a84ecd · outbound

This paper cites Prentice Hall PTR, 1994.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Prentice Hall PTR, 1994

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.294359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.665577Z digest=sha256:87339b7a562510440580f59fe4b317bc7e7520c0a3919cec4dde904fb54e4e41

Observation 22d561b4-8f13-4005-81e5-f64b2b7b4880 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.284042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.668777Z digest=sha256:b03d146fa5c86b6052d690f0a4e23dc3472de1b6a2fe8d17d0e07d66eb46bae4

Observation 7701116d-ee02-4b81-99de-e9b508b57309 · outbound

This paper cites Video activity localisation with uncertainties in temporal boundary.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Video activity localisation with uncertainties in temporal boundary

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.273144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.671586Z digest=sha256:2d1015e46e066478a610109c82362d1d201876137c2d0a670de8a3e23d87f899

Observation 96fe039f-1857-4ac6-aba6-d6b8fe041f4f · outbound

This paper cites Cag-qil: Context-aware actionness grouping via q im- itation learning for online temporal action localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Cag-qil: Context-aware actionness grouping via q im- itation learning for online temporal action localization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.263894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.674278Z digest=sha256:032d3d56b103e05494edf3cc44c71ea4e76fdad6f5e8b86336c51cbc77541afb

Observation b347f31d-2c9f-4030-81bd-a2a602ebd70a · outbound

This paper cites A sliding window scheme for online temporal action localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding A sliding window scheme for online temporal action localization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.252811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.677553Z digest=sha256:79c96f9d69d14ed60217cda010206f614291eff97b7abc466e1b25cf9b456847

Observation a38911d8-ad59-4ccc-9e24-2b367f69981d · outbound

This paper cites Dense-captioning events in videos.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Dense-captioning events in videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.243918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.681058Z digest=sha256:7c28f75e76f112965d6de4243ab788b125440acd5a72113ac4540a2ebe6049ff

Observation 61f5090b-c327-43a5-bd3e-bdccb56cdc8c · outbound

This paper cites Efficient adaptive human-object inter- action detection with concept-guided memory, 2023.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Efficient adaptive human-object inter- action detection with concept-guided memory, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.684597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.684597Z digest=sha256:d39400c5f7cae815c2290dda334b1f06e8835e64fa9847d2cffde5432c573b65

Observation e705e4b5-20a3-4d7d-9d95-b2cb8646400c · outbound

This paper cites Cross modal adaptive few-shot learning based on task depen- dence.Chinese Journal of Electronics, 32(1):85–96, 2023.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Cross modal adaptive few-shot learning based on task depen- dence.Chinese Journal of Electronics, 32(1):85–96, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.228389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.688413Z digest=sha256:27c1eee54c740a7c203429d29c87068f0e0105fd6b891769095382a9153180b6

Observation 416f8863-437c-4ce6-ab09-062868f6a599 · outbound

This paper cites G2l: Semantically aligned and uniform video grounding via geodesic and game theory.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding G2l: Semantically aligned and uniform video grounding via geodesic and game theory

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.218186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.691376Z digest=sha256:6bf73f7e97ceade9a32b9cbd4df2989ae6eec6f353440a685f7e78d70d83c537

Observation 92a0772e-0c2a-4ed5-8f40-ed514094ea8d · outbound

This paper cites Mo- mentdiff: Generative video moment retrieval from random to real.Advances in neural information processing systems, 36, 2024.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Mo- mentdiff: Generative video moment retrieval from random to real.Advances in neural information processing systems, 36, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.208553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.694304Z digest=sha256:6262da7b247f2989a9b5769287896c4e67e37445c376cfc2773de88077120324

Observation ea08fb25-3a90-4b9d-a5e6-d6e13bc7ceb8 · outbound

This paper cites Focal Loss for Dense Object Detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Focal Loss for Dense Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.698296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.698296Z digest=sha256:da3a7dc9b4dd4b01435f56bd008496d5a908a22ffe8efb105859646f74cc5a64

Observation 90fbad3b-e4c0-40fc-897e-d323c11b0485 · outbound

This paper cites Towards balanced alignment: Modal-enhanced semantic modeling for video moment re- trieval.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Towards balanced alignment: Modal-enhanced semantic modeling for video moment re- trieval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.198675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.701578Z digest=sha256:41e8e553b4641423c58342d20c8b833e736f708ae693f9df1bea32bad8697ba9

Observation 586018d7-d057-490a-9fb5-f9c59ebb7acd · outbound

This paper cites Decoupled Weight Decay Regularization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.704250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.704250Z digest=sha256:7c6b2bbb99d88b0a826480e262b6d69c76f9398cf19c3290709909659e5b57a3

Observation bbc6bdd9-1d23-428d-83a0-cc781903a67b · outbound

This paper cites Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.187780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.707413Z digest=sha256:b117637eca724be46a7930414147cfce671fcde238dffe01194d3d16a7514ec9

Observation 682839cf-a0e6-4a5e-bb05-b3cf8e05db07 · outbound

This paper cites Zero-shot video moment retrieval from frozen vision-language models.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Zero-shot video moment retrieval from frozen vision-language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.178388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.709898Z digest=sha256:72758be9acde82c838c987d27502d951534a9b10d558f28180cb07f93e3ef1c8

Observation 1c511538-94b7-4b1d-be38-a696d8d0e853 · outbound

This paper cites Snag: Scalable and accurate video grounding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Snag: Scalable and accurate video grounding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.167028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.712473Z digest=sha256:c9397841aa0ccb0af6648946a1e84feda7f670e2357a900c403c2ef28a781a00

Observation 29c00aae-dbda-4bac-8ab3-666817837939 · outbound

This paper cites Local- global video-text interactions for temporal grounding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Local- global video-text interactions for temporal grounding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.157203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.715246Z digest=sha256:b6bb74dcdcced8c15ba436174bc1993ea4897af61e6681700f060b2409c4c2c3

Observation 62a2d395-520d-4e59-bda1-26624444bab3 · outbound

This paper cites Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.717769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.717769Z digest=sha256:262dbb560db29cac41ac1c92ecacc763498260be2b4c451181e41ad0fcdd2255

Observation 55159b80-fd16-46c9-923d-d042b8b564fe · outbound

This paper cites An overview of cross-media retrieval: Concepts, methodologies, bench- marks, and challenges.IEEE Transactions on Circuits and Systems for Video Technology, 28:2372–2385, 2017.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding An overview of cross-media retrieval: Concepts, methodologies, bench- marks, and challenges.IEEE Transactions on Circuits and Systems for Video Technology, 28:2372–2385, 2017

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.147813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.720559Z digest=sha256:523c5217cd760aa6d37e96e003ca0983beebd91e30f100bfe6f8278ff30fcb97

Observation ecc23ea0-ba85-4e35-965a-b952cf0ef620 · outbound

This paper cites Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems, 37:119336–119360,.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems, 37:119336–119360,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.723182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.723182Z digest=sha256:a60db7917a476a2b38693482b9d5a858a6532e608df9bc75b759233359f84ff2

Observation de278ce3-8708-465c-b82b-22be9f626f49 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.726376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.726376Z digest=sha256:e74fe7d669eb60a7eeaf28dbc44c52c061b6532ac0ecbce00fe40c48a9f77227

Observation ed936847-2d8d-4196-afb2-639f8251bbd4 · outbound

This paper cites Hat: History-augmented anchor transformer for on- line temporal action localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Hat: History-augmented anchor transformer for on- line temporal action localization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.126061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.729869Z digest=sha256:d0cf769937b7a5f19162c24e6318795b620391563acbd9bbe6f885eb973e3b7f

Observation 77504ab3-10f2-404f-9263-8aeda83c81e6 · outbound

This paper cites Coherent multi-sentence video description with variable level of detail.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Coherent multi-sentence video description with variable level of detail

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.116314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.732731Z digest=sha256:6ea6803a797cd542d6e10729ccc2320964f01278776469432f47f80a55704263

Observation d59f1833-1610-415e-a2cd-397bb2018bc6 · outbound

This paper cites Online action detection in untrimmed, streaming videos-modeling and evaluation.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Online action detection in untrimmed, streaming videos-modeling and evaluation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.106866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.735226Z digest=sha256:2ff06492e460ffde9c29bdafb9ed4bfd5d8c47318616230eb494e5da56ff8b4d

Observation 567aa239-1c2b-4965-bc43-e1b8b8bde4f6 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.097528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.737976Z digest=sha256:8c556e791d44ef18f0f77700755c2b3adb43a31e222007a2f7780db3f2e72e9e

Observation 59b7e413-3c40-45c9-9f79-845b88d3bf28 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Moviechat: From dense token to sparse memory for long video understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.087010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.740748Z digest=sha256:c01d3326c1902071d03fcc3e91be2a8ae92769fda3fff39efc780a0a252555cc

Observation e1182f0a-4318-4e2e-ad82-76ff4e3e519e · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Learning spatiotemporal features with 3d convolutional networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.743312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.743312Z digest=sha256:fc60681cfa60ea2740c73ea8fc681aeeb0bd262b7c90ea594f5bfa0840fece2f

Observation a033cea5-7d45-4ca8-b740-712720ff790e · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.072095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.746238Z digest=sha256:6ad15e7b66291b947665bd03358840406d3259a5da4882a0e937f678f8eafd8b

Observation 0e9005ce-0fdf-4021-94d6-b52be4d1e664 · outbound

This paper cites Oadtr: Online action detection with transformers.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Oadtr: Online action detection with transformers

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.062710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.748889Z digest=sha256:84197371c2e80d067da25eabdfd6b9fe44b33e0124396bd18c05a541735a5e12

Observation 6e01456c-dabb-4883-965f-2ae7c1926e44 · outbound

This paper cites Efficient temporal extrapolation of multimodal large language models with temporal grounding bridge.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Efficient temporal extrapolation of multimodal large language models with temporal grounding bridge

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.053554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.751591Z digest=sha256:f48a6e03ea1cdd13afd422d5b7d638d9efb35b1f2b973bc49916e662ef1e40f2

Observation 6c2b59d2-197a-469b-83ea-a207d082543a · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.754369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.754369Z digest=sha256:94fd8c889449aa172877932244997d92ebffbbfaa42fa606b4a87f9aceb6d370

Observation 781b5f8a-ddd4-4500-9748-5d442a38d4f9 · outbound

This paper cites Negative sample matters: A renaissance of met- ric learning for temporal grounding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Negative sample matters: A renaissance of met- ric learning for temporal grounding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.044174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.757493Z digest=sha256:a005ec7503f71f27eb20519f236db209c5e7c9a725891c1441085e2e9c08e71d

Observation 3537a7a4-eb0e-4060-908c-938ad3258cb0 · outbound

This paper cites Longvlm: Efficient long video understand- ing via large language models.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Longvlm: Efficient long video understand- ing via large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.760159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.760159Z digest=sha256:912ebb21e08969b4bbb353f62c6f591d6ae0e116da9c3eba089307e7ce336da7

Observation 70e4c475-a2c6-4257-af6c-805eba6e4563 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.026870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.762870Z digest=sha256:3149328048e9612a7099564c3e1fc9831040070260348e96a9f92d07b073c4e2

Observation 221385eb-fec1-4fb8-a76c-18145febdba6 · outbound

This paper cites Temporal recurrent networks for online action detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Temporal recurrent networks for online action detection

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.765400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.765400Z digest=sha256:5d2714053b81429933e67860ffa89bd7007a736bee80c647ed55b64edfc13dd1

Observation 4d79a847-4f31-4660-a963-303c645526f9 · outbound

This paper cites Long short-term trans- former for online action detection.Advances in Neural In- formation Processing Systems, 34:1086–1099, 2021.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Long short-term trans- former for online action detection.Advances in Neural In- formation Processing Systems, 34:1086–1099, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.011062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.768204Z digest=sha256:e08a2fcb06819ee9e03f7e5bfc640daf8885b3cc0dcea5b1af2f8992a3313ad2

Observation 746f2438-8cdd-40bd-8e19-b901c06f55d0 · outbound

This paper cites Active object detection with knowledge aggregation and distillation from large models.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Active object detection with knowledge aggregation and distillation from large models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.001334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.770695Z digest=sha256:1d168d69e3f56605a2fdc9470cf4bf721de9cba4581a21515a8634da53566334

Observation 6c0b1b3b-d15a-4bca-9b8d-3d113aa5ee20 · outbound

This paper cites Flowgananomaly: Flow-based anomaly network intrusion detection with adver- sarial learning.Chinese Journal of Electronics, 33(1):58–71,.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Flowgananomaly: Flow-based anomaly network intrusion detection with adver- sarial learning.Chinese Journal of Electronics, 33(1):58–71,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.992050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.773233Z digest=sha256:8d8990d7f80020a732018b1b11ba2cb9efc51000dcf532bfd43fb8d2a0fb305c

Observation 2ba91c3e-102d-4de4-a5c1-c99f9b4b88de · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.777349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.777349Z digest=sha256:ac6e2b13a691cb6fd865b8ba1dbdd5f5b4ca8a8b158a323fac5b8a401c14ae9c

Observation 2e867451-fc1f-4ef8-b0fb-a92529a3c4eb · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.780650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.780650Z digest=sha256:f4d4011d70baad557dacef12fd842a310b93c35cb74be55fcb7b06c8ec554fbf

Observation 72ce5440-5286-4afc-aff1-fc9f3b1b4b4a · outbound

This paper cites Progressive privileged knowledge distillation for on- line action detection.Pattern Recognition, 129:108741,.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Progressive privileged knowledge distillation for on- line action detection.Pattern Recognition, 129:108741,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.783553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.783553Z digest=sha256:d7c6e475374efb492924de6a5780085924c2a743077a675809c94e09b2a54f4d

Observation 37c5946a-cc81-4186-8ffd-5bfe6b5459ba · outbound

This paper cites Real-time online video detection with temporal smoothing transformers.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Real-time online video detection with temporal smoothing transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.970000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.786834Z digest=sha256:1739de4405cd146550722495164617255285338f2d918e5ed29dddc75a8061be

Observation 194be584-97ee-4265-a4a1-44f2a41d1b06 · outbound

This paper cites Unsupervised cross-media hashing learning via knowledge graph.Chinese Journal of Electronics, 31(6):1081–1091, 2022.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Unsupervised cross-media hashing learning via knowledge graph.Chinese Journal of Electronics, 31(6):1081–1091, 2022

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.959228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.790152Z digest=sha256:10b9b6d556e8d46595d24f8e644bea28b4a3b4500d121eb1f85144eef070445c

Observation 667feeca-3c4f-479d-a0f3-a69890e413fd · outbound

This paper cites Weakly supervised video moment localization with con- trastive negative sample mining.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Weakly supervised video moment localization with con- trastive negative sample mining

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.947534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.792730Z digest=sha256:fab1b58ed0c7037c0fef3fa72cfdbffb605ec883ce3d26d830b0996a8eb3eb5d

Observation 17cf5ad9-acdc-4bdd-a50b-57c3bc1917ae · outbound

This paper cites Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learn- ing.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learn- ing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.935291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.795687Z digest=sha256:31955ca71b1c4574d323f099011838c6a8c984049be4cebf239cfac8ac346fb7

Observation 74e5305e-d75a-4166-8060-ac84589ef354 · outbound

This paper cites Generating structured pseudo labels for noise- resistant zero-shot video sentence localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Generating structured pseudo labels for noise- resistant zero-shot video sentence localization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.925397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.798721Z digest=sha256:3a20ce57cb5c1e72526ed718cdbfbf6b11e8f7d059327295a72ad3ed547b3d2b

Observation 786ca839-0407-4b2b-abe6-e4301fabdf27 · outbound

This paper cites Phrase-level temporal relationship mining for temporal sentence localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Phrase-level temporal relationship mining for temporal sentence localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.915647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.801411Z digest=sha256:9d33849ec7e1395412b163b3256c3d4413a8cf4477674556ba573b1ab26e145b

Observation f0a2fcde-42ce-43b0-937a-15b51ae2dc2a · outbound

This paper cites Training-free video temporal grounding usinglarge-scale pre-trained models.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Training-free video temporal grounding usinglarge-scale pre-trained models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.905828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T00:00:03.804048Z digest=sha256:e6f00da5dbe850b47ba57767895bb9f68823fc793935d0485b793d0818c4b98b

Observation 9d37da64-71e8-46f9-b5ea-5b159f06b07d · outbound

This paper cites Faster and better learning for bounding box regres- sion., 2020, 34.DOI: https://doi.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Faster and better learning for bounding box regres- sion., 2020, 34.DOI: https://doi

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-08-06T00:00:03.806881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.806881Z digest=sha256:72d19c6891c29cbfb6908ff664cfca3c83506369b642371971a5acc4bba0cd8e

Pith citing papers

No inbound Pith citation observations are available.