Pith. sign in

Paper Citation Record · LEDGER

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2507.21945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21945 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:26.355829Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 293d6fb2-0392-45ce-9586-c84c1a239967 · outbound

This paper cites Finediving: A fine-grained dataset for procedure-aware action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Finediving: A fine-grained dataset for procedure-aware action quality assessment

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:35.144656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:21.844647Z digest=sha256:f9165bb706cd025368f0f16f6b5c733b181419c49499de476eb2fdf8627467db

Observation e9c96844-2699-418e-9bf9-aea2a583ef50 · outbound

This paper cites Fine-grained spatio-temporal parsing network for action quality assess- ment.IEEE Transactions on Image Processing, 32:6386–6400, 2023.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Fine-grained spatio-temporal parsing network for action quality assess- ment.IEEE Transactions on Image Processing, 32:6386–6400, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.898184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:21.954456Z digest=sha256:2d0d8d05b29482055e0b3822ce38110146c79ab2acb40f4fcbe6668589ba888f

Observation bce7b54f-aafb-4f97-b38d-1f2e3ad1d1ca · outbound

This paper cites Learning to score figure skating sport videos.IEEE transactions on circuits and systems for video technology, 30(12):4578– 4590, 2019.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning to score figure skating sport videos.IEEE transactions on circuits and systems for video technology, 30(12):4578– 4590, 2019

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.671458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:22.098946Z digest=sha256:599af9b92ee965e2e0bc8763f0d11fd8b02be020ef7051c89af7ab642b5fc7b9

Observation b9c9d0af-3ffd-43c5-bfa0-60fba00ce5e0 · outbound

This paper cites Hybrid dynamic-static context-aware attention network for action assessment in long videos.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Hybrid dynamic-static context-aware attention network for action assessment in long videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.388543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:22.178255Z digest=sha256:ad82f9aa85dff9be5e8ffed590001ded575451e8f93101f7fa01c5385ff1e867

Observation f703451b-5f11-4694-a114-ee6f3d629405 · outbound

This paper cites End-to-end object detection with transformers.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment End-to-end object detection with transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:22.271513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:22.271513Z digest=sha256:57219ef4cbc2f28559f5c6e3b4fea52c757de257900c3d18664eb64682df04bd

Observation fae1c94c-ff63-46bd-b38d-4890b04186f0 · outbound

This paper cites Likert scoring with grade decoupling for long-term action assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Likert scoring with grade decoupling for long-term action assessment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.116411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:22.358013Z digest=sha256:7d37618dd4ab40f0edef66431968d646891ba1aa06a8d4c26647a8a723c1bb51

Observation 0355ddde-df52-4237-93d8-7814f517d19c · outbound

This paper cites Localization-assisted uncertainty score disentanglement net- work for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Localization-assisted uncertainty score disentanglement net- work for action quality assessment

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.861720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:22.478760Z digest=sha256:3db39bbe7236b2d92c52aeca9e4190e6520ff60cfd337b82f89f1732f1c89c23

Observation 10f02535-4362-44a8-a5d1-6e4330fe6d0a · outbound

This paper cites Skating-mixer: Long-term sport audio- visual modeling with mlps.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Skating-mixer: Long-term sport audio- visual modeling with mlps

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.569918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:22.611670Z digest=sha256:2ba959f8b57661a1c8d8e97d31f0c7ae01d1023804337a6dd457752afaf68217

Observation a137ae47-4f6a-4496-a78a-dbd49b3bc654 · outbound

This paper cites Multimodal action quality assess- ment.IEEE Transactions on Image Processing, 2024.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Multimodal action quality assess- ment.IEEE Transactions on Image Processing, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.296808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:22.702717Z digest=sha256:a64c427eeb171108fbcdaa6f452e38afcd81f6aff2d5cb6ca03c022c3b200e6a

Observation 42c0b6d9-8172-420e-9538-900c4e13dc03 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audio-visual scene analysis with self-supervised multisensory features

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.080761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:22.758132Z digest=sha256:9ae2dc2937c70230b2193a51cccee29a7f6c733cf9ead8df91b2bc3fd64f20bc

Observation 3aa4d9d9-625d-42c8-8ca1-a744569cd76a · outbound

This paper cites Dual attention matching for audio-visual event localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Dual attention matching for audio-visual event localization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.859041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:22.823789Z digest=sha256:d9134b055b9b09587223d8e91cfec834c797aaa28e6ba7b91ceae80119a08f70

Observation c5da39c2-3d3a-4cef-98d2-fadf74989a0b · outbound

This paper cites Egocentric deep multi-channel audio-visual active speaker localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Egocentric deep multi-channel audio-visual active speaker localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.408078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.069552Z digest=sha256:bcbf53b1e325b4c9d8d1dcac02191eda58a62e9a4f4b352082ce058431d297f7

Observation d69bf16c-1e26-4758-83d9-628fe87471fd · outbound

This paper cites Assessing the quality of actions.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Assessing the quality of actions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.134002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.128630Z digest=sha256:9ee24cce6431575a2c6e121c1d9f993f7cfbbafa2040246c65e2dd441939e78e

Observation ee7b2f01-1aec-417b-b3bb-efc9d3d894c1 · outbound

This paper cites Learning to score olympic events.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning to score olympic events

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.939914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.200363Z digest=sha256:eeaecd9855db3614ca07e4fd921c2969d0298a358f37c3539f78d5ac7e2f1368

Observation 6133f46d-31a4-4a7e-b90b-a5545b9c8fb3 · outbound

This paper cites Scoringnet: Learning key fragmentforactionqualityassessmentwithrankinglossinskilledsports.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Scoringnet: Learning key fragmentforactionqualityassessmentwithrankinglossinskilledsports

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.780758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.255859Z digest=sha256:093f78c23ab063643b9354b2773c2fb73e3b29ede2e6e0c5dcadee26d763088b

Observation 2c24ff34-4a45-44be-8408-6740c5944ac1 · outbound

This paper cites S3d: Stacking segmental p3d for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment S3d: Stacking segmental p3d for action quality assessment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.471727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.326548Z digest=sha256:03d5604077a0011a516af29617dad7688b1436b32988103cd10c629772031db4

Observation cbc69f9b-b598-4031-884c-9857ee33c4da · outbound

This paper cites What and how well you performed? a multitask learning approach to action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment What and how well you performed? a multitask learning approach to action quality assessment

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.211318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.377327Z digest=sha256:b2b1237612483aa960ae3bf061a4f4bd8c956d3b64a2653392dc57c1341af59b

Observation 86725ee6-d879-47df-aea2-21bae0051371 · outbound

This paper cites Action quality assess- ment using siamese network-based deep metric learning.IEEE Trans- actions on Circuits and Systems for Video Technology, 31(6):2260–2273, 2020.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Action quality assess- ment using siamese network-based deep metric learning.IEEE Trans- actions on Circuits and Systems for Video Technology, 31(6):2260–2273, 2020

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.927185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.455277Z digest=sha256:5fbce9efcceb806566ce705cd221b7c6b818b6c2b779d7b6fc65b8133e8f819e

Observation 05c56ede-cb32-4a2a-9325-622501bcd876 · outbound

This paper cites Group-aware contrastive regression for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Group-aware contrastive regression for action quality assessment

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.770604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.594595Z digest=sha256:5a977c1ced069568f706d89c0ffecc2afe757801b44cdf408c021baacb778356

Observation 1bf7300b-4013-408e-8638-66c6b208df1e · outbound

This paper cites Tsa-net: Tube self-attention network for action quality assess- ment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Tsa-net: Tube self-attention network for action quality assess- ment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.587767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.704637Z digest=sha256:bb1fbcc6e348669876f35f7405c59e800065a056277994ff9c446b269fd7bc1e

Observation 5b1f2cf4-99be-4007-b2fa-335733409e54 · outbound

This paper cites The pros and cons: Rank-aware temporal attention for skill determination in long videos.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment The pros and cons: Rank-aware temporal attention for skill determination in long videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.370377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.751508Z digest=sha256:f66d4d0f2579ddce931d7416069ac2e450707075db766f4e89b46b4320d29727

Observation ce7576cd-de75-4a86-b20d-eff8782e34aa · outbound

This paper cites Logo: A long-form video dataset for group action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Logo: A long-form video dataset for group action quality assessment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.174716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:23.882275Z digest=sha256:8ed77c9e5c948e675c2cf5801ea97f0bc715ee286f4f95196901eb53f23b3fb7

Observation 54681c95-aeda-4cf9-aca8-04ec0f6c8b98 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audiovisual SlowFast Networks for Video Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:23.939472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:23.939472Z digest=sha256:65db1b3072848156cf1e33924e9e906e5f323b19f47a45a7e9e5f28f64307628

Observation 81734ab5-1ad0-4503-83d4-00709ab74d6b · outbound

This paper cites Listentolook: Actionrecognitionbypreviewingaudio.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Listentolook: Actionrecognitionbypreviewingaudio

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.005897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:24.019897Z digest=sha256:eb8a8f0a4bf51e9c5e6167a26739c7b7d2f339f735959648cd505c08da56434c

Observation 565f5d5d-8e86-4cb2-9ba0-be13bcebb5d9 · outbound

This paper cites Cross- attentional audio-visual fusion for weakly-supervised action localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross- attentional audio-visual fusion for weakly-supervised action localization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.617834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:24.162198Z digest=sha256:e52285ec490505f455c56b0f6596aaa0b81a26387845c4ce1138961e7627bcc8

Observation e52dcf7d-07b4-465c-be09-7cb7d4688d30 · outbound

This paper cites Cross-modal background suppression for audio- visual event localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-modal background suppression for audio- visual event localization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.857858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:24.275209Z digest=sha256:8b6410fceecf0a71fd43dcb20597a3ef54e167a2c4b1ddaf243086b31814c321

Observation b365b3dc-c437-428e-b17f-2b8798d840de · outbound

This paper cites Mmw-aqa: Multimodal in-the-wild dataset for action quality assess- ment.IEEE Access, 2024.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Mmw-aqa: Multimodal in-the-wild dataset for action quality assess- ment.IEEE Access, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.724551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:24.438832Z digest=sha256:937c3509311a44495ca03d41e1571dd2a96b8eb9800ba413b62e185b9f3472aa

Observation 5316e41a-8695-4090-a504-b1b094029013 · outbound

This paper cites Vision-language action knowledge learning for semantic-aware action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Vision-language action knowledge learning for semantic-aware action quality assessment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.608867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:24.554309Z digest=sha256:243184dae89fddcaf858ce531bcd11f216748d306220bd3ad454995e38dd1b6a

Observation 373a5f44-f617-4a56-826b-dcf0f2b3c98a · outbound

This paper cites Learning semantics- guided representations for scoring figure skating.IEEE Transactions on Multimedia, 26:4987–4997, 2023.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning semantics- guided representations for scoring figure skating.IEEE Transactions on Multimedia, 26:4987–4997, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.453029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:24.665443Z digest=sha256:7d113ead18e1e3b794efd43921f293634a11b812ee867d816123d807cac08da4

Observation 6466f5bb-3bea-4602-976b-3fd74bde4804 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal and cross-modal attention for audio-visual zero-shot learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.348648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:24.812348Z digest=sha256:731a899878afd42da458ce4e1d1d0359393b0f5f495b93ef98496033974de67e

Observation 28f464e8-07a7-4ac6-af2c-8f45d0800f4e · outbound

This paper cites Cross-attention is not always needed: Dynamic cross-attention for audio-visual dimensional emotion recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-attention is not always needed: Dynamic cross-attention for audio-visual dimensional emotion recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.201321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:24.981888Z digest=sha256:1192fda118a2cc5508ad79fad0bc3ee7b8435c3239a089aed3b8087d4f5b5c51

Observation dc635a20-9a1d-41c8-b400-95e65f7d9f20 · outbound

This paper cites AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:18:26.702155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:25.094671Z digest=sha256:ec62eaa9452c5bc91e9f05320b5a9d2fd8068eb4b47fc1f237617b2d21220787

Observation 81151f82-1911-418a-9873-2a844699de31 · outbound

This paper cites Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:25.205516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:25.205516Z digest=sha256:69d2a66318462b214e444c1d34792ac7c5978c6fd8e95e46963e93950a9aa388

Observation 9a091c2c-f28a-4b37-a51e-d7c82d18db2c · outbound

This paper cites Temporal alignment networks for long-term video.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal alignment networks for long-term video

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.997529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:25.294446Z digest=sha256:7620494c71e18e08502abea4a00a3fff692f47a57339c7c9cda643e7518575fc

Observation 0cce93b2-4bc4-4ec1-b98c-a11a30d0c358 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning.Advances in neural information processing systems, 35:38032–38045, 2022.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal and cross-modal attention for audio-visual zero-shot learning.Advances in neural information processing systems, 35:38032–38045, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.715831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:25.400298Z digest=sha256:783a223ee1cefb482009c0377cd567c5996a15ba7b4f2ec1c5a7588490808c53

Observation 371a4223-90b5-496a-ad60-888a71e4af2e · outbound

This paper cites Video and accelerometer-based motion analysis for automated surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Video and accelerometer-based motion analysis for automated surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.491535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:25.469481Z digest=sha256:e56a64f7f50f31a8aa6bf48d28275950f7d5a109b904aeb4209a0c6c8739302e

Observation be99137b-cb2b-4394-9f9e-23c44703000e · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audio set: An ontology and human-labeled dataset for audio events

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.040623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:25.735323Z digest=sha256:46206fffb530d5c0a6e673f8f46478ade46560f3b3329372d018dc9700846339

Observation cae90976-fcc6-48fb-9880-545d7827f2bc · outbound

This paper cites Action quality assessment with temporal parsing transformer.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Action quality assessment with temporal parsing transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.837087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:25.826751Z digest=sha256:41358e4ea130142d6b90c5bdfa8aaf279effe719a43a02f6a74169b4ef887e74

Observation 47c93c07-361c-486a-bf55-ee1634b7f01a · outbound

This paper cites Learning spatiotemporal features with 3d convolu- tional networks.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning spatiotemporal features with 3d convolu- tional networks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.690611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:25.887329Z digest=sha256:965020529f326971a537813a1df3575c7fe48170901e222f4aca3c121652fd54

Observation a7d2812a-e2a4-4ec1-bf60-f120a2df2321 · outbound

This paper cites Video and accelerometer-based motion analysis for automated 43 surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Video and accelerometer-based motion analysis for automated 43 surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.509449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:25.948741Z digest=sha256:4200641d7703338dc6629d8ce39e7981971c133af06490facbc332e224767ed5

Observation be607b59-d4b6-4031-b12e-9f75cf801bc5 · outbound

This paper cites Deep resid- uallearningforimagerecognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Deep resid- uallearningforimagerecognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.296707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:26.008721Z digest=sha256:4bcd197ef8605bcf324812d088a08025e61e2b2ed5e07c2175d60e77b03d8e91

Observation b63b5445-ccb8-4bc5-b274-386495198d89 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Quo vadis, action recognition? a new model and the kinetics dataset

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.289591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:26.071958Z digest=sha256:9208e5a03dc25a00fc32c1f1ab41b35738fcdb996b8b94ed4adb6e8e0a6ac045

Observation 187f7f20-0af3-4e6f-a632-27b512db3adc · outbound

This paper cites AST: Audio Spectrogram Transformer.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment AST: Audio Spectrogram Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:26.116606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:26.116606Z digest=sha256:56187f19b0c27138c8fffb4d5c342cdd0c9c5aa878af46c55ca40e305b14fca3

Observation 7f0a990d-0458-4ace-99a2-e20b3a350c74 · outbound

This paper cites Joint visual and audio learning for video highlight detection.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Joint visual and audio learning for video highlight detection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.086584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:26.176463Z digest=sha256:c3988e8a08d4a04e788dead38a07242e90bdb911c74097a1cf3a97c892fcad36

Observation da49cd3c-3a8c-444d-a127-ce5d3ba5b860 · outbound

This paper cites MSAF: Multimodal Split Attention Fusion.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment MSAF: Multimodal Split Attention Fusion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:26.255413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:26.255413Z digest=sha256:7dd5262834e7702e253dbc69b9e5215c479a30e7b2348e407d2b74e1e1fdc2a0

Observation be1ce44f-b478-4363-9b88-7c5c7f62d816 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:26.885434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:26.355829Z digest=sha256:58c6d1dcc9b49f22820a5bc7894faee8e96ff7ae0817201a5b8694fca5bd0b31

Pith citing papers

No inbound Pith citation observations are available.