Pith. sign in

Paper Citation Record · LEDGER

USV: Towards Understanding the User-generated Short-form Videos

As of 11 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2605.20838.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20838 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T04:52:12.045880Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact18
  • verified fuzzy55
  • unresolved6
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff876afc-6a36-40ab-9b36-9af33a36be37 · outbound

This paper cites com / JaidedAI / EasyOCR.

USV: Towards Understanding the User-generated Short-form Videos com / JaidedAI / EasyOCR

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.544592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:5535b86c63695395619ac928629f6a4aad3cd3d2926a587b9d99eacc014c9d64

Observation 901cd8c9-c185-498f-a3d5-69a8103f57ab · outbound

This paper cites an unresolved cited work.

USV: Towards Understanding the User-generated Short-form Videos Unresolved cited work

Reference 2

Resolution
parse uncertain
raw_fallback, observed 2026-05-21T04:53:58.554433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:2857fa5cf7d9d3689948f3d6175582b27743e67e4b7f8ff011a39201af722b91

Observation b705764f-edd8-4dba-a7b1-ec896af76d48 · outbound

This paper cites an unresolved cited work.

USV: Towards Understanding the User-generated Short-form Videos Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:53:58.537448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:0aff8287a1d4c6694f7657b02f731904d5fdb93963d5a8dc9b01abee91461afc

Observation 16e33f42-856c-4b75-b7fc-4d212791d74b · outbound

This paper cites an unresolved cited work.

USV: Towards Understanding the User-generated Short-form Videos Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:53:58.559443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:1b6a6e770db0373af74277ae454283660856c1aa5757ee105eb7411e5f13674d

Observation c3af6144-195c-4050-9a78-346a4e033744 · outbound

This paper cites an unresolved cited work.

USV: Towards Understanding the User-generated Short-form Videos Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:53:58.540984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:d91f60711fba415f32a51d31a65c1809ce9ce1a83194e652421a8e6ad704dcd6

Observation 644911d4-2fef-401d-ae34-c236d945a385 · outbound

This paper cites com / sloria / TextBlob.

USV: Towards Understanding the User-generated Short-form Videos com / sloria / TextBlob

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.586712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:fbc8ce16770b8eaa24d28f1c6c74829eb7ba9495e376e37c1f6e49f59fc904da

Observation 5aec84b2-4bc6-4034-868c-06651d03e95b · outbound

This paper cites an unresolved cited work.

USV: Towards Understanding the User-generated Short-form Videos Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:53:58.588531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:5288cec65f8990ea985d4353bab1259b996712e82b684deed7631f123f4b882d

Observation 4e9b55ba-cb25-45dd-8d89-19f70ca7a017 · outbound

This paper cites an unresolved cited work.

USV: Towards Understanding the User-generated Short-form Videos Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:53:58.584822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:0c12913dae49e183cea58299de0ca98fd86cdf0030ae06d0513a666c9c72da4a

Observation 15731782-7fdf-4608-847b-b6c49668ba02 · outbound

This paper cites businessofapps.

USV: Towards Understanding the User-generated Short-form Videos businessofapps

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.581523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:da9071664ecac4defe8c7014747a63e75829bf0a9b79d6f6743f3333e5099a08

Observation 8519ed7c-b374-48da-a38e-ef49ffd1b072 · outbound

This paper cites YouTube-8M: A Large-Scale Video Classification Benchmark.

USV: Towards Understanding the User-generated Short-form Videos YouTube-8M: A Large-Scale Video Classification Benchmark

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.808325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:394c52397a4bfbedc8fc7cce165c834c7ce2f99696c66eaa9a70c6e5cb215d52

Observation 6626b449-136c-4bd7-aad5-f7d53897b651 · outbound

This paper cites Localizing mo- ments in video with natural language.

USV: Towards Understanding the User-generated Short-form Videos Localizing mo- ments in video with natural language

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.583201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:6aa85ba2c6480bccacd2e4f62fa366333ebff796224c434f49627a1e3df47477

Observation 40c0c428-198b-44d1-b8dc-26651fd8ada0 · outbound

This paper cites Look, listen and learn.

USV: Towards Understanding the User-generated Short-form Videos Look, listen and learn

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.590301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:7a626d4b8a92098df99e2db5652dbace22ebaaf5c8cfe860cefa579a33f6b25c

Observation 203898c1-cb93-4315-9785-dffe800451e5 · outbound

This paper cites A Short Note about Kinetics-600.

USV: Towards Understanding the User-generated Short-form Videos A Short Note about Kinetics-600

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.830162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:8ce87d8e18bf38d6688e163126f6b35a25233f6af73c42d03c203433e58abf7d

Observation e9e93f9e-7bcf-45e0-a799-474b0798142b · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

USV: Towards Understanding the User-generated Short-form Videos Quo vadis, action recognition? a new model and the kinetics dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.594722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:211e70bb66de7500b93272d8492b50039c0e0718d64996b755c7ef089b5fb003

Observation 0182f6ed-f1de-4490-bbdc-f9c3fd075739 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

USV: Towards Understanding the User-generated Short-form Videos Scaling egocentric vision: The epic-kitchens dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.614855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:d2feff7db93befcad19e0516a7271f6630bfd71f84910c3d96d5c00e8e5b41d1

Observation 20cbedf4-95c9-4956-b705-ac3541108f85 · outbound

This paper cites The epic-kitchens dataset: Collection, challenges and base- lines.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11):4125–4141.

USV: Towards Understanding the User-generated Short-form Videos The epic-kitchens dataset: Collection, challenges and base- lines.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11):4125–4141

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.607842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:61def7a14923771ad8388ca20d12f62ef5f6767a4dc7b5519cef6492dc46c84d

Observation 026d33e8-03fa-45aa-a196-c5ba545f2629 · outbound

This paper cites The youtube video recommendation system.

USV: Towards Understanding the User-generated Short-form Videos The youtube video recommendation system

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.643047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:2fd38384d9ded9ea551cc4cb555714481e942217657119fc29b01567b28ff17e

Observation ce4e25bd-20c2-4d95-b511-9f9747c5ae23 · outbound

This paper cites an unresolved cited work.

USV: Towards Understanding the User-generated Short-form Videos Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-21T04:53:58.637695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:d86d93c833ddbf95ffb877a77bb52316eaa0c529fb9d021e71a9b7df079b51a7

Observation 6c9624c9-7429-490a-8d6d-155d148cc9b1 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

USV: Towards Understanding the User-generated Short-form Videos BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.832574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:2222af01f816081108b1693ed70c06d17bc2d786780c00e2c34773a76010913c

Observation 88e0a802-fd8e-4678-8c44-43c5884d7ff9 · outbound

This paper cites Large scale holistic video understanding.

USV: Towards Understanding the User-generated Short-form Videos Large scale holistic video understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.634233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:844bbdac34ca3340054c0a66d06e0854ba709dd3183efa46d7a60191d2685567

Observation 816ca43c-6b58-44a0-8f3c-93e3f67b0272 · outbound

This paper cites Large Scale Holistic Video Understanding.

USV: Towards Understanding the User-generated Short-form Videos Large Scale Holistic Video Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.825012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:fbd950a0c3463c45468f1495b458db77196d9c1c93b73141f4f9ca87665496e0

Observation 79949599-e9c2-4169-8717-35902693c6df · outbound

This paper cites Pyslowfast.https://github.

USV: Towards Understanding the User-generated Short-form Videos Pyslowfast.https://github

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.635892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:01d314658c7ab1598a721ef8d876f1a87669d054fc9458676849979819fbf033

Observation ae50963c-cd8c-49db-a614-05838d34a89f · outbound

This paper cites Slowfast networks for video recognition.

USV: Towards Understanding the User-generated Short-form Videos Slowfast networks for video recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.639336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:143d230a6c1b046c473ec448482f2436298485c7e17bcd25bda07046e7ffad87

Observation e2c013f4-511e-4452-b898-e81114a94569 · outbound

This paper cites Self-supervised video representation learn- ing with odd-one-out networks.

USV: Towards Understanding the User-generated Short-form Videos Self-supervised video representation learn- ing with odd-one-out networks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.626497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:a58504fbbebd2819415c6d1bcf0eb6d725225e9647f66f9ff247b267a0948c2b

Observation 5d6791fb-72a0-4885-81c0-c2bf22b8f07e · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

USV: Towards Understanding the User-generated Short-form Videos The” something something” video database for learning and evaluating visual common sense

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.628501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:d64e467ec04881034f2a1b02e6d72d6f0203f534879928a24ceec34108f8c06b

Observation 92dc1cd7-1a68-40f5-ab06-2ff374c53145 · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions.

USV: Towards Understanding the User-generated Short-form Videos Ava: A video dataset of spatio-temporally localized atomic visual actions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.630242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:1981e5ffe45341ef0f8a845d2533214422f805face7f3ac340312f2c56bf9904

Observation 5356342a-f475-4391-8ca9-2d97748288f5 · outbound

This paper cites Video rep- resentation learning by dense predictive coding.

USV: Towards Understanding the User-generated Short-form Videos Video rep- resentation learning by dense predictive coding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.620549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:847744261ae2ae1c2ad4e9d0a0827e3e35b666196966416fa881b887fbd1ecfd

Observation 8fa54744-ddd6-4de7-b7ae-36378d79d66b · outbound

This paper cites Deep residual learning for image recognition.

USV: Towards Understanding the User-generated Short-form Videos Deep residual learning for image recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.622746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:e2a4df9988134eec08f9b5a397ae089f5a4513ca113482d85c61c01cf9636a63

Observation 5a56fc33-4973-46b4-8c8c-c466bff290c5 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 961–970.

USV: Towards Understanding the User-generated Short-form Videos Activitynet: A large-scale video benchmark for human activity understanding.2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 961–970

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.642707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:35cdb97bb66ed6314f7786e76f6f01cf436398da7d65f74b3c327482401dc130

Observation 4da275d2-7af1-4596-8af0-1bb98553466b · outbound

This paper cites A hierarchical deep temporal model for group activity recognition.

USV: Towards Understanding the User-generated Short-form Videos A hierarchical deep temporal model for group activity recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.656369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:d2e4a22a61828d19c0d885cabdb52d731b1483ba8a92009fd550614a8f1d95e1

Observation 082ba956-6261-4c00-8d21-06df14aa8059 · outbound

This paper cites Query- aware sparse coding for web multi-video summarization.In- formation Sciences, 478:152–166.

USV: Towards Understanding the User-generated Short-form Videos Query- aware sparse coding for web multi-video summarization.In- formation Sciences, 478:152–166

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.654428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:ac7a9d471cf9cf4fe2cafc91ac8fcad207e6d26d2b4ceaa51eca16dd07a7d0b8

Observation 76cf20e9-68a4-49ae-9fd2-6072ca26f326 · outbound

This paper cites Thumos challenge: Action recognition with a large number of classes.

USV: Towards Understanding the User-generated Short-form Videos Thumos challenge: Action recognition with a large number of classes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.624549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:89491c5eec08b1a0a29d72bffc2355b48e7f6c7b75b7da1e666914260d581e32

Observation 86a0ac54-197b-4e4d-8ada-9737ae8e14c8 · outbound

This paper cites Large-scale video classification with convolutional neural networks.

USV: Towards Understanding the User-generated Short-form Videos Large-scale video classification with convolutional neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.632354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:74d018a869e80a0a92335cc77117ce90cfb1e4c8eb6d704249b54ce492dda2c0

Observation 89e27a36-1452-4a36-bb6b-e707e6f54cee · outbound

This paper cites Large-scale video classification with convolutional neural networks.

USV: Towards Understanding the User-generated Short-form Videos Large-scale video classification with convolutional neural networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.640985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:2b5dc4b045af1fc13cd7bda573fb4f230d7d193138dd369070cd22f4c9d8b906

Observation 9eea8bdc-159c-458f-9646-fdb207889cbc · outbound

This paper cites The Kinetics Human Action Video Dataset.

USV: Towards Understanding the User-generated Short-form Videos The Kinetics Human Action Video Dataset

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.796179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:996489d0be5149153b56eb19cf1d52cbf18c2a7302d4bb18dffab9f1ea3ebc34

Observation 0e1b290d-e503-4de6-9a14-707e777fb2fa · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

USV: Towards Understanding the User-generated Short-form Videos Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.842984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:ea26602df857ee1a67339b1cbc7776e7d54d2065c35d1db2b85e39103a0d1969

Observation 228610bb-69c2-4f3f-b093-bb62e44ef332 · outbound

This paper cites Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization.

USV: Towards Understanding the User-generated Short-form Videos Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.822385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:a853d0e72527380ec396a1e7de83e3c89fabd69a75d1e9fb99ef408fbcc74780

Observation 4c65784f-f784-4d81-ade0-b6df543c8355 · outbound

This paper cites Dense-captioning events in videos.

USV: Towards Understanding the User-generated Short-form Videos Dense-captioning events in videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.602063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:03fb679ba94c4930c710d47fd1645ff2525447f49ece2db8ffe71818dc512a2d

Observation 0840cfe8-9337-4ed4-bcaa-18e37452d9c4 · outbound

This paper cites Hmdb: a large video database for human motion recognition.

USV: Towards Understanding the User-generated Short-form Videos Hmdb: a large video database for human motion recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.603822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:b15c998b81d909006e1720ba435ebf633f6f6c5c1ced361ce91d794471d4e8fa

Observation 76a1ac55-9124-4e31-a77d-2e8011fffc0b · outbound

This paper cites Unsupervised representation learning by sort- ing sequences.

USV: Towards Understanding the User-generated Short-form Videos Unsupervised representation learning by sort- ing sequences

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.600362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:093e88213c0e0f555074f9a1d507d7354475d7f53251c72c24603e6516fe0812

Observation 51298f50-9feb-46a9-8d0c-d5981141c788 · outbound

This paper cites Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling.

USV: Towards Understanding the User-generated Short-form Videos Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.840651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:db5c5838e72865fc6629fc3fb26818eca065f3fd733da6fa0b5903dd5f46b643

Observation 41f40409-73ee-4aed-b6f5-d977c26c61c3 · outbound

This paper cites Learning Spatiotemporal Features via Video and Text Pair Discrimination.

USV: Towards Understanding the User-generated Short-form Videos Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.838074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:c65612a8c6eb80c03b9576a1b9c299c41a045451756e6a7169bf0f80d7fd4ce5

Observation 9c815742-de14-45cd-b6b4-bb0bb8e40f61 · outbound

This paper cites Visual semantic search: Retrieving videos via complex tex- tual queries.

USV: Towards Understanding the User-generated Short-form Videos Visual semantic search: Retrieving videos via complex tex- tual queries

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.605972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:06eae3e74fb7ec5b43270dcd15cb988972ffbc640fc7d9d087663b906e5d4987

Observation df6edc6d-bc24-4b7c-afbd-0e590442bc7e · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

USV: Towards Understanding the User-generated Short-form Videos Tsm: Temporal shift module for efficient video understanding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.609640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:9d5db56353e4935d3a4a53d3c58e00c87b9f38d5a8d9201eefd5020144827467

Observation 03308ac7-ee4f-4d74-9333-e54be3360139 · outbound

This paper cites PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding.

USV: Towards Understanding the User-generated Short-form Videos PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.817082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:144389246739b7bba9085fccd98b91905e3294bc9172c08ae7365e6f2e578225

Observation b73bc4f7-da0a-48bd-beac-833459a6c27a · outbound

This paper cites Towards micro-video understanding by joint sequential- sparse modeling.

USV: Towards Understanding the User-generated Short-form Videos Towards micro-video understanding by joint sequential- sparse modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.596526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:fc11ffe9af10240113a0b6db024d5e2c66c7c2949294dc65ce74f6713572c86b

Observation 03e7c293-58ff-4918-8f19-a81bcc62bba1 · outbound

This paper cites Visualiz- ing data using t-sne.Journal of machine learning research, 9(Nov):2579–2605.

USV: Towards Understanding the User-generated Short-form Videos Visualiz- ing data using t-sne.Journal of machine learning research, 9(Nov):2579–2605

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.598432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:8a3af0eb4be45d325e516bd50e5c6550c581be3e5aed595e33a1e40f479dfc1c

Observation 5cfa8307-338c-43b2-86a4-c15a5d20a7a3 · outbound

This paper cites The jester dataset: A large-scale video dataset of human gestures.

USV: Towards Understanding the User-generated Short-form Videos The jester dataset: A large-scale video dataset of human gestures

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.611294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:e9351c7c0c40b7fef743773a4c30de827004098e0b62751ae1a18055c2dfcced

Observation 4e068227-8c32-4e40-a63f-a7b2fd45aeea · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

USV: Towards Understanding the User-generated Short-form Videos End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.576185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:0513d9cc29811c330336542e9f1e143ad35992e3ff9e28a2136a87ab5c19d23f

Observation 04316950-5876-422d-acfb-ed2ea017f449 · outbound

This paper cites Learning a Text-Video Embedding from Incomplete and Heterogeneous Data.

USV: Towards Understanding the User-generated Short-form Videos Learning a Text-Video Embedding from Incomplete and Heterogeneous Data

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.814545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:935e63d20021e6b8b783169981207a052d8c7e847d3fffbe379159d06cd92f21

Observation 56d4ce16-050d-4ab4-9b06-adbc66bd7b55 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

USV: Towards Understanding the User-generated Short-form Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.578008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:0521562aeaed26bca3544dfe4119f6be4269173480ee1de4b76311dc06f930ba

Observation b86f6cf5-f17d-4e65-8e83-f77930905373 · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

USV: Towards Understanding the User-generated Short-form Videos Moments in time dataset: one million videos for event understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.579742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:27c0689252390c2bbc5896d2039f93bf47bc13a8b3523083488efd912a426472

Observation a65bc5b0-5909-4e00-aa76-30d159434846 · outbound

This paper cites Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding.

USV: Towards Understanding the User-generated Short-form Videos Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.809493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:257d04cab3015d3889ef20ea79410522bcfce0fff7c519397534048b2212bf3a

Observation d80d571c-151a-4484-9f47-0b13bdc92fb4 · outbound

This paper cites Multimodal learning toward micro-video understanding.Synthesis Lec- tures on Image, Video, and Multimedia Processing, 9(4):1– 186.

USV: Towards Understanding the User-generated Short-form Videos Multimodal learning toward micro-video understanding.Synthesis Lec- tures on Image, Video, and Multimedia Processing, 9(4):1– 186

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.571701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:5efe958f16c6a7c8d0b003dcaf8ecd66fa27c634d7cd12861392a6ead054fd07

Observation 651867aa-767a-430c-9a9f-0f731bb9dbdd · outbound

This paper cites Enhancing micro-video understanding by harnessing external sounds.

USV: Towards Understanding the User-generated Short-form Videos Enhancing micro-video understanding by harnessing external sounds

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.594990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:1df9d97cd194a7b2aa28ce1ce5b6e1985ca4677fea29ef821ae09ecafa99536b

Observation 5775011b-aab5-4dd5-9019-c0fb76849cb4 · outbound

This paper cites A large- scale benchmark dataset for event recognition in surveillance video.

USV: Towards Understanding the User-generated Short-form Videos A large- scale benchmark dataset for event recognition in surveillance video

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.652131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:42926690743858947ebd013a08830b03ed03a7f9c479a8a4d17aa3116d7eec10

Observation 3dd1dc4b-e8b2-41a4-8f09-9eb5c93856b9 · outbound

This paper cites Learning joint representations of videos and sentences with web image search.

USV: Towards Understanding the User-generated Short-form Videos Learning joint representations of videos and sentences with web image search

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.650245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:57a8e3a09aace5a76889166c69eb912f44ee4c36d8cfb2e181c50d7ed881cf81

Observation fd6c52a3-9073-481e-9dbc-537898075de3 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

USV: Towards Understanding the User-generated Short-form Videos Learning Transferable Visual Models From Natural Language Supervision

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.806444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:a2c15317ffc522d91854a4c6cebf036943088d58259f256b3f2eb80120296c88

Observation b9a22372-02c1-4e89-a37b-63c5c65a9383 · outbound

This paper cites Scenes-objects- actions: A multi-task, multi-label video dataset.

USV: Towards Understanding the User-generated Short-form Videos Scenes-objects- actions: A multi-task, multi-label video dataset

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.644724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:92e3b4a9cdd96a31d7969684ad8844ae5c372fdf5d274d2c085eb065a76619dd

Observation d3fbbf07-9dea-4fae-b6e6-869e98cd9cdb · outbound

This paper cites A dataset for movie description.

USV: Towards Understanding the User-generated Short-form Videos A dataset for movie description

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.646399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:2c27763916d15744218a23039a93cf7eb3065ef89c6f71818f3631169b2bc0db

Observation b8e92a28-c075-4055-ad6e-e7d518d18d86 · outbound

This paper cites Movie description.International Journal of Computer Vision, 123(1):94–120.

USV: Towards Understanding the User-generated Short-form Videos Movie description.International Journal of Computer Vision, 123(1):94–120

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.638862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:090375828dcabe85ec67a190733e79f290dfb77a1dd0a8b4370c834e76eaa9f6

Observation 75295bc4-abb6-4897-90ec-68792ebec72a · outbound

This paper cites Ntu rgb+ d: A large scale dataset for 3d human activity anal- ysis.

USV: Towards Understanding the User-generated Short-form Videos Ntu rgb+ d: A large scale dataset for 3d human activity anal- ysis

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.640730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:2f406f3842427ffcbe79e0607e87bb62cfc269e72a01877d55f90dd5e87b6410

Observation 52f66faf-fce0-45d8-b18f-4c7cf013f027 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understand- ing.

USV: Towards Understanding the User-generated Short-form Videos Finegym: A hierarchical video dataset for fine-grained action understand- ing

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.647083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:7ec01feaf5fed22f8b1e22913f8452f95e211e7938cb90eed9ea78889e7c0c92

Observation 5df7220e-6658-4b02-8d90-02f339d8fd3e · outbound

This paper cites Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision.

USV: Towards Understanding the User-generated Short-form Videos Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.819712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:9b92bb47b22f0ac2c7ed3cdb0a1dce76f1df872dfb25116430e83bb24aff267e

Observation 04450844-ca60-447e-bd35-2ba4e203b489 · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

USV: Towards Understanding the User-generated Short-form Videos Two-stream con- volutional networks for action recognition in videos

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.634997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:54e99f264afa76edcfcc720c23efd6b8d25d99c02de6ed4847b39718c59bb3db

Observation 80fff4f4-3a66-43b4-9f5b-9c3b74af7d78 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

USV: Towards Understanding the User-generated Short-form Videos UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.811877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:fa070bd49cd4dfb0b02237fc2ad91e5b87376fb46c7dd34074c44b743353b6e1

Observation 8072604e-d5ae-4f54-ba1d-09a61eb1cdea · outbound

This paper cites Learning Video Representations using Contrastive Bidirectional Transformer.

USV: Towards Understanding the User-generated Short-form Videos Learning Video Representations using Contrastive Bidirectional Transformer

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.803596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:0df6f84dd8f03728b3310e04ca54c0ad5bad2e242209d5948a7c1f1907ac3c41

Observation b7c0138d-73a3-4998-890d-9548560e18b4 · outbound

This paper cites Learning Language-Visual Embedding for Movie Understanding with Natural-Language.

USV: Towards Understanding the User-generated Short-form Videos Learning Language-Visual Embedding for Movie Understanding with Natural-Language

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:57.835508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:a1df267a54280364b15878ca20706318c026db9256e84a764226c969defc9674

Observation ff7c9ead-7288-4b54-9348-38317db12868 · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition.

USV: Towards Understanding the User-generated Short-form Videos A closer look at spatiotemporal convolutions for action recognition

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.648140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:02cd07c6b05e4bab18ffa2420b6e4287c9cf04c33fac6be37903d585a17a27ba

Observation 389ff030-b348-40ae-a58d-8b0d69b86823 · outbound

This paper cites Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics.

USV: Towards Understanding the User-generated Short-form Videos Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.631265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:88d5ec46b1be355ee4162c80ce68eb0550b3f032926921f5cba9326baa6a1140

Observation 5ac333e1-0497-4d85-b378-1594fbe60725 · outbound

This paper cites Temporal segment net- works: Towards good practices for deep action recognition.

USV: Towards Understanding the User-generated Short-form Videos Temporal segment net- works: Towards good practices for deep action recognition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.627503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:1c203f5950a51cbdffa974b8d6437e57980bca1b5514db5b21fa35e7352903a8

Observation dcac19e1-98e6-4886-97be-0216cfd00c47 · outbound

This paper cites Neural multimodal co- operative learning toward micro-video understanding.IEEE Transactions on Image Processing, 29:1–14.

USV: Towards Understanding the User-generated Short-form Videos Neural multimodal co- operative learning toward micro-video understanding.IEEE Transactions on Image Processing, 29:1–14

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.625574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:a4edc2b3114df2fa2edd044d4c76718c58ba361adb877df8880d305c3bfbbe7a

Observation 4809551d-1056-40b8-b724-475f4aee7105 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

USV: Towards Understanding the User-generated Short-form Videos Audiovisual SlowFast Networks for Video Recognition

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.827901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:21b566dc224fd65bd1f83d73cef9a061b35f219d7d0d08c1d4174d2881d61a14

Observation 667d61a2-845a-4d0e-85ca-5a6633b23c2d · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

USV: Towards Understanding the User-generated Short-form Videos Msr-vtt: A large video description dataset for bridging video and language

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.629257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:80046b5d8b3602fff9333ab26d3f1698cc7851756574e2962a162e0164383acb

Observation e466b55e-62a4-4a2b-acba-4179377c47c0 · outbound

This paper cites Large-scale weakly supervised audio classifi- cation using gated convolutional neural network.

USV: Towards Understanding the User-generated Short-form Videos Large-scale weakly supervised audio classifi- cation using gated convolutional neural network

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.633120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:b361240b37f9da42af8ff3a833360623b60b721e544206ecd5cab4f9df9622c0

Observation 2e8bea90-9194-4f5f-a1e5-b6a81620f6cc · outbound

This paper cites A joint se- quence fusion model for video question answering and re- trieval.

USV: Towards Understanding the User-generated Short-form Videos A joint se- quence fusion model for video question answering and re- trieval

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.621501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:32677c129e7fce8a330ed20f3a612ebc68673005039e68248a4b5b4fb9f311ac

Observation 87c6bb0a-c7ad-4431-9ff8-b66b21ddad1a · outbound

This paper cites End-to-end concept word detection for video caption- ing, retrieval, and question answering.

USV: Towards Understanding the User-generated Short-form Videos End-to-end concept word detection for video caption- ing, retrieval, and question answering

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.619548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:1fb1a5f775863d99ac2567f7866ab8f12d54f1384961b77902e43a26987894ed

Observation 3bb22a96-317d-4357-bb02-409880b73cab · outbound

This paper cites Low-rank regularized multimodal representation for micro-video event detection.IEEE Access, 8:87266–87274.

USV: Towards Understanding the User-generated Short-form Videos Low-rank regularized multimodal representation for micro-video event detection.IEEE Access, 8:87266–87274

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.617205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:b83bb7c54de2b4a36cbdc74b8291a101c5c373b315682b7e60045bbef40269ff

Observation 4b9d62c7-73e9-4a9d-894b-a74c7f7a74af · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

USV: Towards Understanding the User-generated Short-form Videos Towards automatic learning of procedures from web instructional videos

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.623593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:2aaa6174a440c9cfbf27ed98752b52aedfcde9bdc26158663f8a938e01a3ce62

Observation b4eb4319-34b4-443e-b17e-00fe44866fbd · outbound

This paper cites Videotopic: Content-based video recommendation using a topic model.

USV: Towards Understanding the User-generated Short-form Videos Videotopic: Content-based video recommendation using a topic model

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T04:53:58.615353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:5362d74bdbd08ca09e136e3c379e0f47e6bdef34b0e8e53ddf2c8b1a58637b50

Pith citing papers

No inbound Pith citation observations are available.