Pith. sign in

Paper Citation Record · LEDGER

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning

As of 19 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2505.19938.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19938 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:09.912975Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy63
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56943cd2-bb7d-40db-8ce6-ad0560ce3393 · outbound

This paper cites Contrastive masked autoencoders are stronger vision learn- ers,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Contrastive masked autoencoders are stronger vision learn- ers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.663955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:01.568706Z digest=sha256:436520cd9121489fda906e792f4b110599dee9d4493ab5b9874c7c889a318ae2

Observation 2b640d07-4292-42d4-ace2-7d59051f7b50 · outbound

This paper cites Pyramid constrained self-attention network for fast video salient object detection,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Pyramid constrained self-attention network for fast video salient object detection,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.434976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:01.651719Z digest=sha256:9281d77dcdc98b63e6e5c34f7ed6ac6a81b813965bfa0228cffa9c5d125e2499

Observation 8ef2a54b-4207-4fd3-88b8-a0049b3c53d4 · outbound

This paper cites Cap4video: What can auxiliary captions do for text-video retrieval?.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Cap4video: What can auxiliary captions do for text-video retrieval?

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.199773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:01.737481Z digest=sha256:584d5361855988928ad003195db6ed921fb73c8b67b9b06baa639d719e331e4f

Observation 81e4713e-a90c-4f92-9e21-5d16d6c5f391 · outbound

This paper cites Isomer: Isomerous transformer for zero-shot video object segmentation,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Isomer: Isomerous transformer for zero-shot video object segmentation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.939461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:01.827885Z digest=sha256:1f0deaab4be6e77c385bf6758c7272dd417b828f1f59924962138fece897b7c3

Observation 9774c393-8a97-4a22-bcca-4e2128547b1e · outbound

This paper cites Avgzslnet: Audio-visual generalized zero-shot learning by reconstructing label fea- tures from multi-modal embeddings,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Avgzslnet: Audio-visual generalized zero-shot learning by reconstructing label fea- tures from multi-modal embeddings,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.716658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:01.895004Z digest=sha256:49d5478364b7bc8cad1ad1acc90327d9fa17452ac1197ce68b4bc2950b12fa91

Observation 96dfe273-d3b7-44f3-b65f-168020578767 · outbound

This paper cites Audio-visual generalised zero-shot learning with cross-modal attention and language,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Audio-visual generalised zero-shot learning with cross-modal attention and language,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.495744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:01.947941Z digest=sha256:9f3cfb830003d4c219859cc5781188f59b539c1e77103ba44292cc89a176abd7

Observation 5374bc77-da0d-4a67-99d4-261c17547950 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Temporal and cross-modal attention for audio-visual zero-shot learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.327195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.034750Z digest=sha256:e7abd3c92167f2b0bf637d77967e680093ce645ce060a75c9f2bb6de3a2472c7

Observation 72eadc1a-7ed2-45af-b8bf-4ac01f2ebc22 · outbound

This paper cites Enhancing unsupervised video representation learning by decoupling the scene and the motion,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Enhancing unsupervised video representation learning by decoupling the scene and the motion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.141166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.099314Z digest=sha256:13bcab25f8ddc219f9bc248a75080f4b411e8b35583e568ef40b0302bd802acc

Observation d294de1c-d33d-434d-8345-dfd0b1d232df · outbound

This paper cites Spiking tucker fusion trans- former for audio-visual zero-shot learning,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Spiking tucker fusion trans- former for audio-visual zero-shot learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.897923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.171589Z digest=sha256:855665dcd704a49d328f90fb25252183ca140747416a8ab3a5f53b0b4e531b51

Observation d22828d8-7513-4d2c-9f2f-ea4dd1827329 · outbound

This paper cites Motion- decoupled spiking transformer for audio-visual zero-shot learning,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Motion- decoupled spiking transformer for audio-visual zero-shot learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.698806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.280202Z digest=sha256:55e83bb44a57be05ea1c1f8740af0c67927b94ca17116c7f54f9e57c127f8bd7

Observation cd4268d7-8fc9-4232-b1dc-3848f2550c48 · outbound

This paper cites Conv2former: A simple transformer-style convnet for visual recognition,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Conv2former: A simple transformer-style convnet for visual recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.478377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.344994Z digest=sha256:41a114a0581e79ef4a71fd9aa3a32687c1bc866ced079821c496651b84fb3735

Observation 8d4b895f-2d50-49f8-8933-91b75b429e17 · outbound

This paper cites Bidi- rectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Bidi- rectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.246538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.452965Z digest=sha256:f1abd27ced40a6419a6a7968c77a5aeb544cbf2712256209cb5b045caa97eb12

Observation 152c6262-d221-4388-88d4-ec2ce45aa4ef · outbound

This paper cites Box2mask: Box-supervised instance segmentation via level-set evolution,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Box2mask: Box-supervised instance segmentation via level-set evolution,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.011981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.565696Z digest=sha256:ae82eb1be41272fe30f67a3994480b28ebac353944db30120f1c3782c94c50e3

Observation bc6f5344-5385-4392-8b9c-c5f25606c707 · outbound

This paper cites The devil is in the crack orientation: A new perspective for crack detection,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning The devil is in the crack orientation: A new perspective for crack detection,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:21.768417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.614753Z digest=sha256:1d501f5493888c507a680b20599d3b9e7441da608815a67e0a3618325fbf1fde

Observation 9c838a6c-1fa3-437c-b871-9888bd701059 · outbound

This paper cites Offline and online optical flow enhancement for deep video compression,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Offline and online optical flow enhancement for deep video compression,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:21.562105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.713981Z digest=sha256:b7e8718bd9230c9b8aa2ff7a347b0a71ec4220e45c530caca519a201e33ef96b

Observation e0d569d4-dc6e-496f-90bf-fb5eca6c6ede · outbound

This paper cites Ustc-td: A test dataset and benchmark for image and video coding in 2020s,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Ustc-td: A test dataset and benchmark for image and video coding in 2020s,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:21.383618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.799059Z digest=sha256:216b3756ba6ecb60f2e1cd5fcafafa7f24c0653840665b16e71d42ce257e200a

Observation 9a8775c8-91b0-4118-b3b2-a4ddc6f22524 · outbound

This paper cites Object segmentation- assisted inter prediction for versatile video coding,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Object segmentation- assisted inter prediction for versatile video coding,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:21.214751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:02.886398Z digest=sha256:6b6326affb6cec987874b8fd1c8618fe25aec8cb4b03319a04348b8eb8d8e761

Observation bc236cc0-2058-43a1-a884-2c063075f2b1 · outbound

This paper cites Geometry-aware guided loss for deep crack recognition,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Geometry-aware guided loss for deep crack recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:20.894270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:03.005736Z digest=sha256:fc4a677c1d5fd9d35343bb36fa6b7ecb3b6fa2a12c69253092c615287fce5356

Observation 335a8d7e-b2a8-4b77-a304-5a584cda06ca · outbound

This paper cites Latent embedding feedback and discriminative features for zero-shot classifi- cation,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Latent embedding feedback and discriminative features for zero-shot classifi- cation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:20.667962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:03.178906Z digest=sha256:8d3a357b15f1352e2740f3422ed02472ee87cb6ef943b9df75b4ba56fa7a60f6

Observation aeac26fd-b60c-43a0-85e7-f85d70e3ffdf · outbound

This paper cites Gener- alized zero-and few-shot learning via aligned variational autoencoders,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Gener- alized zero-and few-shot learning via aligned variational autoencoders,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:20.422107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:03.261550Z digest=sha256:8ee3deeaa125ab738cdd5631b411edfcb82cb47334050937c101a632eeadf1f1

Observation 4890ef3a-e3fc-492f-a6f5-431316bf26d1 · outbound

This paper cites Generalized zero- shot learning via synthesized examples,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Generalized zero- shot learning via synthesized examples,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:20.273737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:03.347665Z digest=sha256:31bfeaf18669dc81f70c2f79e6714faa92c91ca4a57c5701ac381442ca242bbb

Observation fa02b7cb-7af7-4bb7-a696-d85af6a46754 · outbound

This paper cites Feature generating networks for zero-shot learning,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Feature generating networks for zero-shot learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:20.056647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:03.436163Z digest=sha256:4a244a7eef0ba024f587be6c3e718557e93922d01414d2cef843002169d9869e

Observation 96beda2a-9a91-4ed1-a624-f49bfb6cec95 · outbound

This paper cites A gener- ative adversarial approach for zero-shot learning from noisy texts,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning A gener- ative adversarial approach for zero-shot learning from noisy texts,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:19.795041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:03.656519Z digest=sha256:576ca7a15c39d50747c47511d42867a511d6be777c90d5328f3c1db2446afdde

Observation 8eb83658-a615-48bd-863b-b75ee70eff8f · outbound

This paper cites Bootstrapping audio-visual video segmentation by strengthening audio cues,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Bootstrapping audio-visual video segmentation by strengthening audio cues,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:19.515338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:03.748830Z digest=sha256:2ec2b2da18ed8f20a3c063fc21981f1135bf01be60f75c892c268518e2c4a8e7

Observation 32f96bad-46e4-402f-9eb1-b8924ee198ce · outbound

This paper cites Learning affective features with a hybrid deep model for audio–visual emotion recognition,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Learning affective features with a hybrid deep model for audio–visual emotion recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:19.308470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:03.884901Z digest=sha256:95cc7bd2b1ad1e531dc492c5c24e868cdcafc14005337cf40f04ac04a7357f85

Observation e0feff2f-b320-4bc7-8f0e-7ddc95a4d5b4 · outbound

This paper cites Audio-visual temporal forgery de- tection using embedding-level fusion and multi-dimensional contrastive loss,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Audio-visual temporal forgery de- tection using embedding-level fusion and multi-dimensional contrastive loss,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:04.015637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:04.015637Z digest=sha256:c5ff6fe48e1fd9e1cdf324ae62fb7968f708124154da1d3e1482f54d05e48cff

Observation 5f40ea82-1016-437e-9738-3f96262091e9 · outbound

This paper cites Multimodal imbalance-aware gradient modulation for weakly-supervised audio-visual video parsing,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Multimodal imbalance-aware gradient modulation for weakly-supervised audio-visual video parsing,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:19.061134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:04.124896Z digest=sha256:579063567e4d88a1c605c4166eaeeebe573cba58fb605d839f2d838c223f4610

Observation 99c8f7ca-4ecb-45b9-9098-c0024099dac6 · outbound

This paper cites Question-aware global-local video understanding network for audio-visual question answering,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Question-aware global-local video understanding network for audio-visual question answering,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:18.816815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:04.244947Z digest=sha256:cc0fd1b7f2c240b6f37c66f35a09fa617c18bad53626d354b26ca1019ecdab4d

Observation 8598ebb1-6c3c-4604-bb54-bdf61f09a116 · outbound

This paper cites Zero-shot audio classification via semantic embeddings,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Zero-shot audio classification via semantic embeddings,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:18.571251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:04.384768Z digest=sha256:aeda48f5873316500d412ff5bbd9e761df0e568d7acba685e68d0727ce956b03

Observation ca30d387-47f6-4d08-bffd-b47bf6915d08 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity under- standing,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Activitynet: A large-scale video benchmark for human activity under- standing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:18.328939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:04.517362Z digest=sha256:095b86115505f8bcdd4aff80d563f4fdfe81bbc87704cf8ff132cffb1e3b7734

Observation b763a99f-a635-4215-876f-5ec52e37bf42 · outbound

This paper cites The multivehicle stereo event camera dataset: An event camera dataset for 3d perception,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning The multivehicle stereo event camera dataset: An event camera dataset for 3d perception,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:04.655427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:04.655427Z digest=sha256:b1992ba514505656451ecf25a275c856ba2f673daefe6cc3c6427c4f7a6447e7

Observation b1d3db10-a6c6-4a30-92b2-24bf19a37cb6 · outbound

This paper cites Events-to-video: Bringing modern computer vision to event cameras,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Events-to-video: Bringing modern computer vision to event cameras,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:18.032441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:04.774982Z digest=sha256:23f67ea5a103ecdf903ac9bb5f759961d509988c575b786ec991ad53ca32ee78

Observation ce3f04ee-9d93-4e6b-b5f4-d1e7d154b55d · outbound

This paper cites High speed and high dynamic range video with an event cam- era,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning High speed and high dynamic range video with an event cam- era,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:17.734144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:04.914751Z digest=sha256:b61c15e6e6fa7bf58eddeb619c2b69ac498694f5f800768588c59949683446cc

Observation 168ebb16-0ac6-4d13-b09d-39f8574077ec · outbound

This paper cites A controlled-delay event camera framework for on-line robotics,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning A controlled-delay event camera framework for on-line robotics,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:17.503876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:05.065292Z digest=sha256:fb1de9e9cb5d11ac6743a4f264709a2d36feb4d214ee8dd91497538f48e49e2f

Observation 790bfafb-2384-473b-919f-fa664d7e7f34 · outbound

This paper cites Exploring event camera-based odome- try for planetary robots,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Exploring event camera-based odome- try for planetary robots,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:17.337725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:05.199442Z digest=sha256:20bef669859d25b9b18f5d5353568801294283ace443ce987874c66f9be4e510

Observation 2bfae8a2-9561-414e-830a-3820729395b6 · outbound

This paper cites Emergent visual sensors for autonomous vehicles,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Emergent visual sensors for autonomous vehicles,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:17.062862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:05.554856Z digest=sha256:2fa672c702e2985ca9f4eb8c153f53412c4d4fd1bf739a43efe5684b5cd01ab6

Observation a201dff8-bc2f-4aa4-9654-00b954b1b25b · outbound

This paper cites Dsec: A stereo event camera dataset for driving scenarios,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Dsec: A stereo event camera dataset for driving scenarios,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:05.686946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:05.686946Z digest=sha256:6e14f69115339b2d8e5bdb3cd267447d375dc2ec7a90bbb65046154e95e7b8f4

Observation 04cb1f09-c62d-4ee6-bed6-1eaab6adbf21 · outbound

This paper cites High frame rate video reconstruction based on an event camera,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning High frame rate video reconstruction based on an event camera,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:05.796504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:05.796504Z digest=sha256:4ff000f9a7fad3ea184475699de926d9123754c2d8d9904b55b4e34bb1ed81a0

Observation 0abe025d-b50e-45ab-997e-152fb4cffb0b · outbound

This paper cites Eventcap: Monocular 3d capture of high-speed human motions using an event camera,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Eventcap: Monocular 3d capture of high-speed human motions using an event camera,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:16.829097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:05.886800Z digest=sha256:f0b7de961410b238e8606c48b2d09ce6cd6d1dccaaaf06a4d1b5b3fd8412b2ed

Observation ca3e4702-39f3-48f1-a5e6-f5a483e1bc4d · outbound

This paper cites Towards a framework for end-to-end control of a simulated vehicle with spiking neural networks,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Towards a framework for end-to-end control of a simulated vehicle with spiking neural networks,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:16.580114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:05.992309Z digest=sha256:f0f4d2109a8627df9e3cc69ab1d32a46cd613d01cdc83e7c722c48d8e4bc27eb

Observation 8ad00811-e3d4-4fd9-9bbe-b2fcd6bb35fd · outbound

This paper cites Esim: an open event camera simulator,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Esim: an open event camera simulator,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:16.297738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:06.104578Z digest=sha256:8073766278ff6111bd9cdca00093d25e854b4a433efee4762a976f43eb87544d

Observation 50984320-2288-4ec0-bdfd-0b7b540096be · outbound

This paper cites Deep residual learning in spiking neural networks,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Deep residual learning in spiking neural networks,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:16.030530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:06.254621Z digest=sha256:5a9941810885c7482c4f24e5dbc4de1cf6930cc3aaf310fc18127c3ba6ae5ae7

Observation 1908f1e8-a23c-43e1-8215-0962b04c3ad5 · outbound

This paper cites Neuron-based spiking transmission and reasoning network for robust image-text retrieval,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Neuron-based spiking transmission and reasoning network for robust image-text retrieval,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:15.810887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:06.375431Z digest=sha256:1452a30fb97a0cdd2006e35909ec17af791f4dcf9b3456ba153110218883e749

Observation 8ea97e88-dfc0-404d-bb79-fc02000280f3 · outbound

This paper cites Spikemba: Multi-modal spiking saliency mamba for temporal video grounding,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Spikemba: Multi-modal spiking saliency mamba for temporal video grounding,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:15.659258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:06.524922Z digest=sha256:947da2369495d079103be96ae703f138614bca494b49a56174a9364dc8f11246

Observation 46043712-dc94-4421-871e-5663e2582dac · outbound

This paper cites Spikformer: When Spiking Neural Network Meets Transformer.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Spikformer: When Spiking Neural Network Meets Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:06.676151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:06.676151Z digest=sha256:92a2f379865d9e1d52cae8e36c1b039a99aae781dd50e8ad89203437657eca28

Observation 6ba149dc-de50-4140-a96c-a68f0935e6ce · outbound

This paper cites Learning optical flow from continuous spike streams,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Learning optical flow from continuous spike streams,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:15.530931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:06.803713Z digest=sha256:bdc063dab997f237375a5de83f26241edc15c90591c46be2c612393b358be05c

Observation 66f12683-fef4-46c6-a877-3248c9c04f75 · outbound

This paper cites Mrdflow: Unsupervised optical flow estimation network with multi-scale recurrent decoder,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Mrdflow: Unsupervised optical flow estimation network with multi-scale recurrent decoder,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:15.259577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:06.924771Z digest=sha256:a20d41dd7b5f209090002210357d98490bf473614393dce459a0ef6cd37134e0

Observation 278be33e-297f-4b06-8070-7a744574ee54 · outbound

This paper cites Progressive tandem learning for pattern recognition with deep spiking neural networks,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Progressive tandem learning for pattern recognition with deep spiking neural networks,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:07.001775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:07.001775Z digest=sha256:f6f11bd8f9f6fe08101085573daa3e46dc2887b28a2fc0bd041d02085ff6c1d6

Observation 8fcc8209-16b4-474f-b5b7-d1365f29a074 · outbound

This paper cites A hybrid neural coding approach for pattern recognition with spiking neural networks,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning A hybrid neural coding approach for pattern recognition with spiking neural networks,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:15.059970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:07.144661Z digest=sha256:e39488ccaa0f3489d21808e93feb41ccd6dde9fcbc6cd68b5d8f3ebac46eb58b

Observation db5e6eba-a867-4b5f-9c9b-9b7ae219c9e0 · outbound

This paper cites Neuron-based spiking transmission and reasoning network for robust image-text retrieval,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Neuron-based spiking transmission and reasoning network for robust image-text retrieval,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:14.787790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:07.244932Z digest=sha256:3c212a8df3311d9972cd5d820856fb26c330c2a845aa1be7fbb624bd6cc6fb2c

Observation 6fd95f30-7122-4b75-a031-d4ac0e3c045b · outbound

This paper cites Multi-scale spiking pyramid wireless communication framework for food recognition,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Multi-scale spiking pyramid wireless communication framework for food recognition,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:14.550685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:07.374851Z digest=sha256:c9b18e0083d8feee5d1e527a047fa8e53b45504d5f51763df75226a9d2ddc5a0

Observation afecb359-12b1-45c5-a368-c58306b4cdd9 · outbound

This paper cites Modality-fusion spiking transformer network for audio-visual zero-shot learning,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Modality-fusion spiking transformer network for audio-visual zero-shot learning,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:14.299387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:07.528788Z digest=sha256:3343ccaeb30eec6de1539120bbd2c8b992b87208accf7509d7c887005e046fcc

Observation c347e535-9f45-417b-96d3-32f7a3f910ff · outbound

This paper cites Evolving spiking neural network controllers for autonomous robots,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Evolving spiking neural network controllers for autonomous robots,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:14.087173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:07.624761Z digest=sha256:d6d20900dffa939393eea309407c1827b349922aac859314d779a9a6d7e9ce90

Observation d514cb12-ea7f-4377-af39-8414dbc4b048 · outbound

This paper cites A hybrid rein- forcement learning approach with a spiking actor network for efficient robotic arm target reaching,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning A hybrid rein- forcement learning approach with a spiking actor network for efficient robotic arm target reaching,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:13.898839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:07.764899Z digest=sha256:e0b8f343666cbb646a8ad2255e6ffd79096e2ace09aed04a3d04772b95ce0f3c

Observation 6197bdc3-8273-4f63-9980-0717138746e5 · outbound

This paper cites Labelling unla- belled videos from scratch with multi-modal self-supervision,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Labelling unla- belled videos from scratch with multi-modal self-supervision,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:13.667100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:07.892721Z digest=sha256:0252a88073fc66026be99acaa27083077b16fbb4271d16fa7fec3c519e1214fe

Observation 62c2a68e-2068-42ef-95c3-d5731f50b2c0 · outbound

This paper cites Training deep spiking neural networks using backpropagation,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Training deep spiking neural networks using backpropagation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:13.309184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:08.014840Z digest=sha256:94218c28aeb7ddb2ee579ca02972c9ece4e6f4d9b98911469cb9109dd3a2d3ba

Observation 0a7f348c-7b08-4a95-879d-6c975b739301 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:08.114827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:08.114827Z digest=sha256:c9007265c57b90fe632f4f40c598c8a3a8330cf22fd0bdefeb200c91b980c15f

Observation 73e029cc-5e64-408c-ba90-87b5c3841221 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Vggsound: A large-scale audio-visual dataset,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:13.063599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:08.294831Z digest=sha256:ad063c22a5522b7684c04c169476aa23a01161acaad83f6116f84e063ebfb9b1

Observation c9ebbbad-7867-4301-93db-4ea49adc30c6 · outbound

This paper cites Evaluation of out- put embeddings for fine-grained image classification,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Evaluation of out- put embeddings for fine-grained image classification,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:12.827084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:08.552122Z digest=sha256:7c5b18580cf4084c4bbd465e5fa679f5448ab9fc641c962361ae98fc3446a825

Observation 2817b972-4b79-4cc2-9ccf-408ce2df92f9 · outbound

This paper cites Devise: A deep visual-semantic embedding model,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Devise: A deep visual-semantic embedding model,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:12.592774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:08.696128Z digest=sha256:7ae5ab700da6951b1c33ec7b58256bdb149795d53535fd297adb572cda85fcfd

Observation f3da65f5-da96-4ccb-813c-e0b2ff9074dd · outbound

This paper cites Attribute proto- type network for zero-shot learning,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Attribute proto- type network for zero-shot learning,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:12.381999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:08.864855Z digest=sha256:571d0ace29ec6aec3f46b597f696396aa25ae2995c715748fe46fe2ba444ef40

Observation 5c73589e-bba8-44ac-92fd-0a93f7fb45cc · outbound

This paper cites f-vaegan-d2: A feature generating framework for any-shot learning,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning f-vaegan-d2: A feature generating framework for any-shot learning,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:12.053420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.054759Z digest=sha256:3fdab706f1d6467fd3a132d55476a3773aca8188d37238c7dda0f9a8b29acafd

Observation 592f091a-326b-4d8e-95ba-2544a116faf8 · outbound

This paper cites Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classifi- cation and retrieval of videos,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classifi- cation and retrieval of videos,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.779300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.224775Z digest=sha256:20b6904b35f4b49e2b898f648399224a59f64a74fb6e7c196f0e228a3b16f6b3

Observation 9d9254a5-804d-4161-9e73-79ecb95ecf07 · outbound

This paper cites Hyperbolic audio-visual zero-shot learning,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Hyperbolic audio-visual zero-shot learning,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.520841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.287939Z digest=sha256:e32c91841770dbff18e12a0b9c4235fddafb7d7aecb63f86698f29eef67cd40c

Observation eb881181-73c0-4686-95a0-5754709a04f3 · outbound

This paper cites Label-embedding for image classification,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Label-embedding for image classification,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.290505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.395070Z digest=sha256:47ab6563193411371a6fd4cade1f282637a46c678f9914338d1afaaee9fed396

Observation e60af83a-2dac-4304-a94f-50d39e6e2620 · outbound

This paper cites Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classifi- cation and retrieval of videos,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classifi- cation and retrieval of videos,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.127256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.495795Z digest=sha256:d70192fb17adef5971b67beddaa4917bedccddc684a421c5ff78815c56beaee0

Observation 13fd456c-432e-403a-8ef0-0484e6d60668 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Learning spatiotemporal features with 3d convolutional networks,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:09.568720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:09.568720Z digest=sha256:71569844af65a3b2889539138051b5b312fb7fecbac63170e1334a24f56b9246

Observation 961118c7-1ec5-4ad6-a11e-f40eea78e88b · outbound

This paper cites Large-scale video classification with convolutional neural networks,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Large-scale video classification with convolutional neural networks,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:10.887396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.628189Z digest=sha256:2bb87f9e23770b894b6fc7d915fb815b2e5159cbbb3da330903c114e2150c6b6

Observation cc69ac4b-8d18-4d93-8c3b-98839b187297 · outbound

This paper cites Cnn architectures for large-scale audio classification,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Cnn architectures for large-scale audio classification,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:10.668104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.714162Z digest=sha256:fb75f2f64048605a69342fca855ad54ef9169223913b5ed423ebd21691b6a011

Observation 7a6f9a88-c432-455e-912e-86a7ffa889ae · outbound

This paper cites Youtube-8m: A large-scale video classi- fication benchmark,.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning Youtube-8m: A large-scale video classi- fication benchmark,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:10.440823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.800570Z digest=sha256:373a98af59e3982704da248426f3ad568f8389792d02f1a7d6b2f37df94a6d22

Observation c083aee1-8488-41d6-ac1f-46f9eb635c3d · outbound

This paper cites degree with the School of Computer Science, Harbin Institute of Technology, Harbin, China.

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning degree with the School of Computer Science, Harbin Institute of Technology, Harbin, China

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:10.197158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:09:09.912975Z digest=sha256:c7f357dbf5d1dba60c3c071c3c19385bb72125988efe36a68d6bf0f0e1a62fc7

Pith citing papers

No inbound Pith citation observations are available.