Pith. sign in

Paper Citation Record · LEDGER

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2506.01466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01466 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:47:49.917130Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T01:03:08.765350Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T01:03:18.406504Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3e61b2b-19ba-46b9-a1e5-e4b51bfae013 · outbound

This paper cites Ub- normal: New benchmark for supervised open-set video anomaly detection.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Ub- normal: New benchmark for supervised open-set video anomaly detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.820160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.257744Z digest=sha256:0fab0f42b12f5483d0f54a99c8133c8a9a308c08c33780b8193769b2abbc36a5

Observation 76be9730-b41b-4d14-be18-93689015a9e8 · outbound

This paper cites Robust real-time unusual event detection using mul- tiple fixed-location monitors.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Robust real-time unusual event detection using mul- tiple fixed-location monitors

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.793423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.294557Z digest=sha256:2fa5dfe4f0b21628e5186f4b91f6dc746fbbb3a4778a0a572131e5e1b571e33a

Observation 2781f05a-17ec-47f1-b6ab-a430a9a21d66 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.769933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.338099Z digest=sha256:368e88e521c1b9879c71e3ae1ff932415602ac956521cb7c08f08a906164d7c1

Observation 24441f28-bd6c-4da7-a083-ff4ce0db8b4d · outbound

This paper cites Context recov- ery and knowledge retrieval: A novel two-stream framework for video anomaly detection.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Context recov- ery and knowledge retrieval: A novel two-stream framework for video anomaly detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.741986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.381728Z digest=sha256:fc401f8b707b5379b58fc1ac2b1c32c76f2ad083ec50e2e1607f5e6bc093afba

Observation d3aa3363-97d0-4a86-91dd-064478ddfa72 · outbound

This paper cites Gramian multimodal representation learning and alignment.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Gramian multimodal representation learning and alignment

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.719392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.417217Z digest=sha256:207b83f0fa8750e5fb085a9555b678f37f3fdae85ccd4e14e53c1d991ebc3a87

Observation a496dd08-4744-4d82-96c5-c2e5ead58977 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.694784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.454077Z digest=sha256:f505810a03dd2fe9c26e1438289b5f54d24bcec02d1fa555678e07f7d2db37e9

Observation f391045a-4b07-402f-8ff7-9b7885b04901 · outbound

This paper cites https://github.com/modelscope/diffsynth- studio, 2023.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark https://github.com/modelscope/diffsynth- studio, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.663266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.487703Z digest=sha256:1b2639420af96f974dd1250dc3d5f7aafadf4521ff7e3ff3c9420d650c134576

Observation aa5dc797-0c68-4faf-869b-de024f8171de · outbound

This paper cites Oops! pre- dicting unintentional action in video.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Oops! pre- dicting unintentional action in video

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.629283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.523469Z digest=sha256:fb0a668688deebe935046a6768b5e96500ca1224bc0f7df0ba3b47d07e748b4d

Observation c8d7672a-e83c-48cf-afef-44107b200500 · outbound

This paper cites Mist: Multiple instance self-training framework for video anomaly detection.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Mist: Multiple instance self-training framework for video anomaly detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.607665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.557768Z digest=sha256:9b9a661670c6c1f76da428e8f34bade28d23a57c5ebeb63f5670861a35d80a4b

Observation 7b00f05f-9335-4bed-81be-05184ab6f5dc · outbound

This paper cites Cnvid-3.5 m: Build, fil- ter, and pre-train the large-scale public chinese video-text dataset.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Cnvid-3.5 m: Build, fil- ter, and pre-train the large-scale public chinese video-text dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.572148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.603164Z digest=sha256:262047f1939099b98537559b8988d8a08aa8f173b7bbfbee954ff4715f33daef

Observation d0861689-6476-4af3-ace4-432f4c20299f · outbound

This paper cites Tem- poral tessellation: A unified approach for video analysis.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Tem- poral tessellation: A unified approach for video analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.534192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.649305Z digest=sha256:70c85608937eb3636a6f49b76cb2115ccaae176aa79cf8cee28f60c78fa87eba

Observation 47aaabf1-c5ca-4621-a95e-a4b3924ddc8f · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:59.685597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:59.685597Z digest=sha256:7d31cdcb8877b80c3151afe1cbef423a4091738897b38b7e15459646c8545ba6

Observation f1a1ffcf-2bac-4b9c-8780-410566b31af2 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.506777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.718189Z digest=sha256:6b9b9ba25ba809c47413ae30113eedbf4f3bfd39904248acaa93ece5614402e0

Observation 0218e602-47e0-4e4d-8f5f-01d32fefa3b6 · outbound

This paper cites Selvaraju, Akhilesh Deepak Got- mare, Shafiq Joty, Caiming Xiong, and Steven Hoi.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Selvaraju, Akhilesh Deepak Got- mare, Shafiq Joty, Caiming Xiong, and Steven Hoi

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.477579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.771665Z digest=sha256:5a6aeb0fa1e43b740889c9e290d0830c7174b8b4eb81b00ef08e6fd29cbf8e08

Observation 80cab21c-fdcd-41f3-ad77-2ce7d66efa75 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.447046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.816199Z digest=sha256:b7b7b759e07f043c1ee088fdee80f09ad5ecc6d38c29dc16b66dc4a3e79c9a9c

Observation 26ccb0db-e76b-49a7-af37-77816ab71122 · outbound

This paper cites Anomaly detection and localization in crowded scenes.IEEE transactions on pattern analysis and machine intelligence , 36(1):18–32, 2013.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Anomaly detection and localization in crowded scenes.IEEE transactions on pattern analysis and machine intelligence , 36(1):18–32, 2013

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.416998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.861918Z digest=sha256:5ad67aae98d91fdf6455a54c4d4928c998e6dd50f6388e5af75549bfd6b97c04

Observation e45c6866-5130-4dd1-9a7e-a2ea3b718453 · outbound

This paper cites Fine-grained key-value mem- ory enhanced predictor for video representation learning.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Fine-grained key-value mem- ory enhanced predictor for video representation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.393712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.897479Z digest=sha256:522d4367beed7145a181deb8cde33761301392bfe233e5d8977d8fbd2ccf79a2

Observation ef58b787-bd5e-4cc7-a471-3dc9d2eab273 · outbound

This paper cites Timestep embedding tells: It’s time to cache for video diffusion model.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Timestep embedding tells: It’s time to cache for video diffusion model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.368644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.944689Z digest=sha256:d9f677aeb0cdc4dc4840efca7e76daa7143bbbf761d851aa3471c4699f767ed4

Observation a784cecb-13ac-41e5-bc38-6e3a1bdd2f63 · outbound

This paper cites Ntu rgb+ d 120: A large- scale benchmark for 3d human activity understanding.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Ntu rgb+ d 120: A large- scale benchmark for 3d human activity understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.341893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:46:59.980551Z digest=sha256:e2d99b2f36b0cc48686396c81e9574be7082c320397dd7e4910a1bba3a6d3046

Observation 57088754-9328-46ed-b572-3baaa8057a17 · outbound

This paper cites Fu- ture frame prediction for anomaly detection – a new baseline.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Fu- ture frame prediction for anomaly detection – a new baseline

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.318743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.015552Z digest=sha256:20a143cd1777cf391ae9609a0b37b286459c892eb534e3eacd327df255e7458d

Observation ba8dd758-a087-4f6f-afc2-8e1b6ba58bf7 · outbound

This paper cites Abnormal event detec- tion at 150 fps in matlab.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Abnormal event detec- tion at 150 fps in matlab

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.296958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.065422Z digest=sha256:5fe8ffa7c6a789016a4f3d7caad2b234acd00270aaef8dadcb99ec91a392556a

Observation f88d50f8-8bb0-491f-9bec-e31856423e61 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:48:04.275100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.110324Z digest=sha256:9b1ac5b4796095e78d62d72c37b3b159d243ef152167e8d87e6d762a66ff7d27

Observation cc15ba76-37ad-4474-95ac-737f93687af2 · outbound

This paper cites A revisit of sparse coding based anomaly detection in stacked rnn framework.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark A revisit of sparse coding based anomaly detection in stacked rnn framework

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:52.230451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.147802Z digest=sha256:9dfd1f4c2b5cb6d95dd939f964ce76387617911ece8a0ea1d97b262552805227

Observation ea7a4e13-c862-48ed-a4fb-8b50de34b642 · outbound

This paper cites X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:52.157747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.182569Z digest=sha256:0c32997f5d5ff13a757a355640c4fc8e068fbb6588e344e06dac3979adf2f62e

Observation 479b3a4a-e8c7-4d05-87b8-793d42d18eed · outbound

This paper cites MULDE: Multiscale Log- Density Estimation via Denoising Score Matching for Video Anomaly Detection.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark MULDE: Multiscale Log- Density Estimation via Denoising Score Matching for Video Anomaly Detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:52.043913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.218181Z digest=sha256:1b3df665f11adb25bc64b610ee4387a0cf7ba46d29cfa4f338335e469fbf6fdd

Observation fe935186-786b-4e09-88b5-551c215b8fc9 · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:51.986421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.263702Z digest=sha256:f6ad2e6830ffc881cf3e7e87500e803c356c383fda5140f2313506c86155bde6

Observation 8a176fbd-0fbc-48c2-b9b2-ace4788a6b9d · outbound

This paper cites Learning memory-guided normality for anomaly detection.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Learning memory-guided normality for anomaly detection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:51.907315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.305971Z digest=sha256:534555503d7e94b0bd483a072d44bd3772bd80feadbc60a43c0bb6ee0756e643

Observation 807ee946-4392-42bc-bf45-cc4779bc7854 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Learn- ing transferable visual models from natural language super- vision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:51.836927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.342033Z digest=sha256:a0a7a2182f3fa350934564bcd93fe8a69e63ba05b59cbcaef7bb010b1b1d9eb4

Observation 55e30f53-fc90-4984-8a21-33957f0e8697 · outbound

This paper cites Street scene: A new dataset and evaluation protocol for video anomaly detection.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Street scene: A new dataset and evaluation protocol for video anomaly detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:51.783052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.378013Z digest=sha256:9c58c187f1723307ce7bb21e6a9109d1e285e05b7227e48cd1126c412943bb2e

Observation 7ceab09c-2a3c-4ab3-8fbf-38ea0212d68f · outbound

This paper cites Deep-cascade: Cascading 3d deep neu- ral networks for fast anomaly detection and localization in crowded scenes.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Deep-cascade: Cascading 3d deep neu- ral networks for fast anomaly detection and localization in crowded scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:51.594973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.421107Z digest=sha256:f8b4c1249cd352d85360df4be4e87e68cc4210cc58bab6306c095250fa555fee

Observation 7912519c-064e-4453-a84a-b65ae92d858c · outbound

This paper cites Seed-thinking-v1.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Seed-thinking-v1

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:47:00.457859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:47:00.457859Z digest=sha256:dd68fa20236cea4aa1eab1da4b1e45b2474ecc03fa43d2b45336c47a28d310d8

Observation 822328db-508d-4f27-930d-ffcb4db7374a · outbound

This paper cites Real-world anomaly detection in surveillance videos.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Real-world anomaly detection in surveillance videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:51.431409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.509178Z digest=sha256:3527423cdf93d53d05a614d635491332a06957a8c5f0b5268082f9b04368354a

Observation 2489e4e5-1b8b-4c57-9127-e95d5f1a4b1a · outbound

This paper cites Learning Language-Visual Embedding for Movie Understanding with Natural-Language.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Learning Language-Visual Embedding for Movie Understanding with Natural-Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:47:00.567024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:47:00.567024Z digest=sha256:5079e1f08970e78dd748d2e4b2908984c50d07e06d2d8f6ec8bf29e7cefdb719

Observation 86774dba-c311-4625-9e30-c87efad533db · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Wan: Open and Advanced Large-Scale Video Generative Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:47:00.600380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:47:00.600380Z digest=sha256:ffe98605d089fcf85224b06efb7de60ff6313d427bfbf773c76af3e81c8f0353

Observation d631f271-bb79-44bf-8efc-850dfbeac467 · outbound

This paper cites Align and tell: Boosting text-video re- trieval with local alignment and fine-grained supervision.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Align and tell: Boosting text-video re- trieval with local alignment and fine-grained supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:51.215256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.634336Z digest=sha256:a0d63ea1779fe0ffaa93d8ff2177dc801890f8da45893d46a377011a9d4e2100

Observation 5a9a5dc6-6722-4e3f-b70b-85b0f3e7175e · outbound

This paper cites Rtq: Rethinking video-language under- standing based on image-text model.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Rtq: Rethinking video-language under- standing based on image-text model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:51.043543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.668010Z digest=sha256:1ec5a87c8d625ad68bd4e98ae4cbcc99d2d8b6f95ed7a9569f226f292a1de3f3

Observation a1308361-3255-44f3-9c1f-59b3f487f0aa · outbound

This paper cites Weakly-supervised spatio-temporal anomaly detection in surveillance video.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Weakly-supervised spatio-temporal anomaly detection in surveillance video

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:50.914803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.702477Z digest=sha256:ef5aafe54d40313eec18af69fca2aab49441345487b2cbb3e9194be8534c9940

Observation 56210328-3d42-429e-adde-bc05dbf14b7c · outbound

This paper cites An Empirical Study of Frame Selection for Text-to-Video Retrieval.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark An Empirical Study of Frame Selection for Text-to-Video Retrieval

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:47:50.068685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.746919Z digest=sha256:0d1a8751a2f10afeb9f70b9ae8b176b9b82f703ec094205500128b4abd6421da

Observation 726c3cb4-152b-4f38-b109-c0b94bbe65a7 · outbound

This paper cites A deep one-class neural network for anomalous event detection in complex scenes.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark A deep one-class neural network for anomalous event detection in complex scenes

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:50.847482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:00.789124Z digest=sha256:ab4d13afa655226485d1a85e20a03cf958a876e1233a909a4d77d6ab4e94baaa

Observation 011d23c4-624b-4443-8ccb-d5032e634352 · outbound

This paper cites Not only look, but also listen: Learning multimodal violence detection under weak supervision.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Not only look, but also listen: Learning multimodal violence detection under weak supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:50.736171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:49.031838Z digest=sha256:ffc3c1cfc12a7e14cff51555fdbb9f4ab91929cb619ff794a4746100820d7d55

Observation 0f655683-0618-4411-a9ba-631b06229c93 · outbound

This paper cites Toward video anomaly retrieval from video anomaly detection: New benchmarks and model.IEEE Transactions on Image Processing, 33:2213–2225, 2024.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Toward video anomaly retrieval from video anomaly detection: New benchmarks and model.IEEE Transactions on Image Processing, 33:2213–2225, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:50.646165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:49.593594Z digest=sha256:107cc9c584258e548ccfa29ff9b724de0cf42afc80b8df1df0c11d72b12da6c4

Observation 055c5596-8eba-47ff-951d-4c9d7a07e594 · outbound

This paper cites Qwen3 Technical Report.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Qwen3 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:47:49.662405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:47:49.662405Z digest=sha256:13c869bd8e829cbfb0c7e5ea03deba939de521e3bb118ef41489c94a02ef507c

Observation 7544813d-fc40-408b-951c-f9d34cf2ccc0 · outbound

This paper cites A joint se- quence fusion model for video question answering and re- trieval.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark A joint se- quence fusion model for video question answering and re- trieval

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:50.572592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:49.716576Z digest=sha256:8011cd41e603fa46d604135260b42bcee09f912e5d63003a22ba1f68e9b16841

Observation 85981e0f-2e06-4b70-bfdf-d4ce52478934 · outbound

This paper cites Towards surveillance video-and-language understanding: New dataset baselines and challenges.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Towards surveillance video-and-language understanding: New dataset baselines and challenges

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:50.477398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:49.775107Z digest=sha256:aebf381c9c7090af2597754fc63ad14e17ba594551c67d5a20b92805e6fef918

Observation 13787662-58bb-43ef-9e44-16aea6ba8ebd · outbound

This paper cites Generative cooper- ative learning for unsupervised video anomaly detection.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Generative cooper- ative learning for unsupervised video anomaly detection

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:50.386022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:49.834248Z digest=sha256:b4096654da15fdcfa7d7d8c3e6e3536f077655423368a695d6b582637262b342

Observation 594de4d2-954b-4c8d-bd16-75a52231f0fa · outbound

This paper cites Multi-grained vi- sion language pre-training: Aligning texts with visual con- cepts.

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark Multi-grained vi- sion language pre-training: Aligning texts with visual con- cepts

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:47:50.296830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:47:49.917130Z digest=sha256:f8b578141e02f920ce6a3fed8c2b083f93a69e6a55a93155c6160354d2a32472

Pith citing papers

Observation 506c1b59-eaca-43c7-9567-820ad5617057 · inbound

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training cites this paper.

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:03:18.410531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T01:03:08.765350Z digest=sha256:37008350376c8635f87a505a2b803ee976977ec785ab8087c94f397aee2ed18d