Pith. sign in

Paper Citation Record · LEDGER

Audiovisual SlowFast Networks for Video Recognition

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2001.08740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2001.08740 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.185270Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:57:47.660611Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 89b87024-bfb2-4ced-b430-40c7b94df63b · inbound

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models cites this paper.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Audiovisual SlowFast Networks for Video Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.185270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.185270Z digest=sha256:b9c224e050dbe63076fb5ccfa9c468dc56688481f27e42b585d17f924a22584a

Observation 1e498db1-61bb-4662-9211-c4c727c6e716 · inbound

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision cites this paper.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Audiovisual SlowFast Networks for Video Recognition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:24.733808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:24.733808Z digest=sha256:3f9d729467dc14a308edc9e765516beb119c8e698df54f5f100d09b0ffc2f84e

Observation 0bd3bc9d-c4ef-470e-ac41-5a75624ca8e4 · inbound

DMAF-Net: An Effective Modality Rebalancing Framework for Incomplete Multi-Modal Medical Image Segmentation cites this paper.

DMAF-Net: An Effective Modality Rebalancing Framework for Incomplete Multi-Modal Medical Image Segmentation Audiovisual SlowFast Networks for Video Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:12.311803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:12.311803Z digest=sha256:fe4ae8abcc438165893f54a9bcac52a522e808c6975345ab4adf352438cb3cd8

Observation eeeb6297-1c94-4403-b82a-165461958cee · inbound

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding cites this paper.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audiovisual SlowFast Networks for Video Recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.288755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.288755Z digest=sha256:651a653fcf29428f8553a116cd8f4aa4316f739a8b2e9f50897facabb04d0102

Observation 56600a93-ee73-4a74-9978-01870ce512b2 · inbound

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos cites this paper.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audiovisual SlowFast Networks for Video Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.945121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.945121Z digest=sha256:b7743d0db164948e71edae4e41a025c1d29d5fb6a7a8f3bfd7a6aaee144c321f

Observation d4f48735-2087-4877-8af9-04073b46d0bd · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Audiovisual SlowFast Networks for Video Recognition

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.240129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.240129Z digest=sha256:f56428c462ad2231832de512cb0225f3631fc867718a2314cafb3813233d7d47

Observation 54681c95-aeda-4cf9-aca8-04ec0f6c8b98 · inbound

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment cites this paper.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audiovisual SlowFast Networks for Video Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:23.939472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:23.939472Z digest=sha256:65db1b3072848156cf1e33924e9e906e5f323b19f47a45a7e9e5f28f64307628

Observation 67a76d82-262b-4c59-a43b-2d046a5d4749 · inbound

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection cites this paper.

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection Audiovisual SlowFast Networks for Video Recognition

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:13:16.640679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T21:10:01.391718Z digest=sha256:5256de0050d19f36266aed14ef4148a2325cef101caee06b8daa0843ce67214b

Observation 7cbf2acf-08ef-4009-9fd3-f3db5b8f0b6c · inbound

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization cites this paper.

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization Audiovisual SlowFast Networks for Video Recognition

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.544513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:00:44.575524Z digest=sha256:c8fad5afca3b36bcaa8df4337ea873549d844e29826d2bd8b101941effcf8086

Observation 4809551d-1056-40b8-b724-475f4aee7105 · inbound

USV: Towards Understanding the User-generated Short-form Videos cites this paper.

USV: Towards Understanding the User-generated Short-form Videos Audiovisual SlowFast Networks for Video Recognition

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.827901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:b95b76af439d71aea3f50eb0d9cba7a8dd9e5c60f29de0e46cdb616fed5809f5

Observation d10bf92b-4f89-4da0-bbc0-aa8103feef67 · inbound

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning cites this paper.

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning Audiovisual SlowFast Networks for Video Recognition

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:05:06.125333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T21:55:53.630732Z digest=sha256:1851b740355f7699276f702097e446d8a10342f8af26638997e9ce9971f0ec63

Observation 8a8885ba-b003-423e-a10c-cc72a3821ff8 · inbound

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning cites this paper.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audiovisual SlowFast Networks for Video Recognition

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.661964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:11302aab22efbfdff4bdd659ebb6ab36d8e15712005280da8077cc163032899d