Pith. sign in

Paper Citation Record · LEDGER

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2507.21945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21945 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:26.355829Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 293d6fb2-0392-45ce-9586-c84c1a239967 · outbound

This paper cites Finediving: A fine-grained dataset for procedure-aware action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Finediving: A fine-grained dataset for procedure-aware action quality assessment

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:35.144656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:21.844647Z digest=sha256:3a2fea1f4cb04aef17f1993b93a05d5f63519770e92574f9628eef271e7ed69e

Observation e9c96844-2699-418e-9bf9-aea2a583ef50 · outbound

This paper cites Fine-grained spatio-temporal parsing network for action quality assess- ment.IEEE Transactions on Image Processing, 32:6386–6400, 2023.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Fine-grained spatio-temporal parsing network for action quality assess- ment.IEEE Transactions on Image Processing, 32:6386–6400, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.898184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:21.954456Z digest=sha256:c92f5284ed09ead0148cc6dd81b477ad6101418329e8f50b721ce12b71868ddd

Observation bce7b54f-aafb-4f97-b38d-1f2e3ad1d1ca · outbound

This paper cites Learning to score figure skating sport videos.IEEE transactions on circuits and systems for video technology, 30(12):4578– 4590, 2019.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning to score figure skating sport videos.IEEE transactions on circuits and systems for video technology, 30(12):4578– 4590, 2019

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.671458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:22.098946Z digest=sha256:ad8297a59baf1b7401488b48d25805e9292bf20ad340c47b0c67cd56b16c48e4

Observation b9c9d0af-3ffd-43c5-bfa0-60fba00ce5e0 · outbound

This paper cites Hybrid dynamic-static context-aware attention network for action assessment in long videos.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Hybrid dynamic-static context-aware attention network for action assessment in long videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.388543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:22.178255Z digest=sha256:7882069eddcfcaa448a7868c86981aaa515858f2f39408e9d722e1fd78a15a15

Observation f703451b-5f11-4694-a114-ee6f3d629405 · outbound

This paper cites End-to-end object detection with transformers.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment End-to-end object detection with transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:22.271513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:22.271513Z digest=sha256:57219ef4cbc2f28559f5c6e3b4fea52c757de257900c3d18664eb64682df04bd

Observation fae1c94c-ff63-46bd-b38d-4890b04186f0 · outbound

This paper cites Likert scoring with grade decoupling for long-term action assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Likert scoring with grade decoupling for long-term action assessment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.116411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:22.358013Z digest=sha256:36d3715138a1d95e92d59544bc38699e464c921dee9e39031842f11d5f6eb53d

Observation 0355ddde-df52-4237-93d8-7814f517d19c · outbound

This paper cites Localization-assisted uncertainty score disentanglement net- work for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Localization-assisted uncertainty score disentanglement net- work for action quality assessment

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.861720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:22.478760Z digest=sha256:0341f695aa5e9c1bdc102b49f5c233fbaeb483cf9badc558a3cc7aef918f3a90

Observation 10f02535-4362-44a8-a5d1-6e4330fe6d0a · outbound

This paper cites Skating-mixer: Long-term sport audio- visual modeling with mlps.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Skating-mixer: Long-term sport audio- visual modeling with mlps

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.569918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:22.611670Z digest=sha256:538974b7fd2c6f718514b7dfbb993a2523adc428bc876367485684528abc47a5

Observation a137ae47-4f6a-4496-a78a-dbd49b3bc654 · outbound

This paper cites Multimodal action quality assess- ment.IEEE Transactions on Image Processing, 2024.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Multimodal action quality assess- ment.IEEE Transactions on Image Processing, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.296808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:22.702717Z digest=sha256:eef413055cbfac7c0f68054345758a4eecfb7bcfd801587b182389b5db50566a

Observation 42c0b6d9-8172-420e-9538-900c4e13dc03 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audio-visual scene analysis with self-supervised multisensory features

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.080761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:22.758132Z digest=sha256:e7eb9b36ebbb8213c7fc78aec6050c053fe8a6ccc90e726ac95fd044b4ff3c9d

Observation 3aa4d9d9-625d-42c8-8ca1-a744569cd76a · outbound

This paper cites Dual attention matching for audio-visual event localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Dual attention matching for audio-visual event localization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.859041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:22.823789Z digest=sha256:79c13bb44c7ab53438820aa9f892354230be0aa3a64d27b88a7054c0c8507e4f

Observation c5da39c2-3d3a-4cef-98d2-fadf74989a0b · outbound

This paper cites Egocentric deep multi-channel audio-visual active speaker localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Egocentric deep multi-channel audio-visual active speaker localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.408078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.069552Z digest=sha256:e27c1ca8c9767ff80b6e28e1a6a699f3cb68fa4665053a9512cbb972172f5ec8

Observation d69bf16c-1e26-4758-83d9-628fe87471fd · outbound

This paper cites Assessing the quality of actions.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Assessing the quality of actions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.134002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.128630Z digest=sha256:67c6d104af5f24202c1372a94b1755b3be1400c22791f53bb99c305e9974bdfe

Observation ee7b2f01-1aec-417b-b3bb-efc9d3d894c1 · outbound

This paper cites Learning to score olympic events.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning to score olympic events

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.939914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.200363Z digest=sha256:c460fc506d0cad59766ed89bc9f5a7af572b5bf441a0935eef7206d46b72a029

Observation 6133f46d-31a4-4a7e-b90b-a5545b9c8fb3 · outbound

This paper cites Scoringnet: Learning key fragmentforactionqualityassessmentwithrankinglossinskilledsports.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Scoringnet: Learning key fragmentforactionqualityassessmentwithrankinglossinskilledsports

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.780758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.255859Z digest=sha256:1eb344a1bfea7089a052a22d485327f3300b9606123e68fa02bf83aeffc832b1

Observation 2c24ff34-4a45-44be-8408-6740c5944ac1 · outbound

This paper cites S3d: Stacking segmental p3d for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment S3d: Stacking segmental p3d for action quality assessment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.471727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.326548Z digest=sha256:516535a198f132220efc3ce453deecb61e2de8fa3bd228bbeb3aa2a27ce76050

Observation cbc69f9b-b598-4031-884c-9857ee33c4da · outbound

This paper cites What and how well you performed? a multitask learning approach to action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment What and how well you performed? a multitask learning approach to action quality assessment

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.211318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.377327Z digest=sha256:77ecbe029d00e3f9b8397d70a61efbafa37dde929ce9f01341caf96bd1c26d44

Observation 86725ee6-d879-47df-aea2-21bae0051371 · outbound

This paper cites Action quality assess- ment using siamese network-based deep metric learning.IEEE Trans- actions on Circuits and Systems for Video Technology, 31(6):2260–2273, 2020.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Action quality assess- ment using siamese network-based deep metric learning.IEEE Trans- actions on Circuits and Systems for Video Technology, 31(6):2260–2273, 2020

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.927185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.455277Z digest=sha256:418ff3adfc626059423db49bbd80da264d8291721cc3e4c3a029bcecd4525d1b

Observation 05c56ede-cb32-4a2a-9325-622501bcd876 · outbound

This paper cites Group-aware contrastive regression for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Group-aware contrastive regression for action quality assessment

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.770604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.594595Z digest=sha256:fff9995caff3664a47809829016088594abd5ac707dd20e57ea054e5958ffcbb

Observation 1bf7300b-4013-408e-8638-66c6b208df1e · outbound

This paper cites Tsa-net: Tube self-attention network for action quality assess- ment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Tsa-net: Tube self-attention network for action quality assess- ment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.587767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.704637Z digest=sha256:886495733ad73a96877c2969ff5ad754f82cb7bde7b7f4ab35339d4e2236bded

Observation 5b1f2cf4-99be-4007-b2fa-335733409e54 · outbound

This paper cites The pros and cons: Rank-aware temporal attention for skill determination in long videos.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment The pros and cons: Rank-aware temporal attention for skill determination in long videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.370377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.751508Z digest=sha256:5f6842adbdb8c46903e05707827304012f9b30c4f6e07d400d4f09c5a367494a

Observation ce7576cd-de75-4a86-b20d-eff8782e34aa · outbound

This paper cites Logo: A long-form video dataset for group action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Logo: A long-form video dataset for group action quality assessment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.174716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:23.882275Z digest=sha256:9965cc7d865d70cd606e5829540d1811b047a477908663e489a58726241398b3

Observation 54681c95-aeda-4cf9-aca8-04ec0f6c8b98 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audiovisual SlowFast Networks for Video Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:23.939472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:23.939472Z digest=sha256:65db1b3072848156cf1e33924e9e906e5f323b19f47a45a7e9e5f28f64307628

Observation 81734ab5-1ad0-4503-83d4-00709ab74d6b · outbound

This paper cites Listentolook: Actionrecognitionbypreviewingaudio.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Listentolook: Actionrecognitionbypreviewingaudio

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.005897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:24.019897Z digest=sha256:a910ba49aa5b992fb8bdca6c08829a0a197af519bfe98bfd2e80f022afbe3e4b

Observation 565f5d5d-8e86-4cb2-9ba0-be13bcebb5d9 · outbound

This paper cites Cross- attentional audio-visual fusion for weakly-supervised action localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross- attentional audio-visual fusion for weakly-supervised action localization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.617834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:24.162198Z digest=sha256:e5afcc988d98b4fe309ebdbba148eb1c1d25097ac5e7e7231207df6d9b30e6db

Observation e52dcf7d-07b4-465c-be09-7cb7d4688d30 · outbound

This paper cites Cross-modal background suppression for audio- visual event localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-modal background suppression for audio- visual event localization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.857858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:24.275209Z digest=sha256:e4ad5630ef1841b96662b1c44cf7832d1a6e9b4867e63b14e304b9fd04e9ff8d

Observation b365b3dc-c437-428e-b17f-2b8798d840de · outbound

This paper cites Mmw-aqa: Multimodal in-the-wild dataset for action quality assess- ment.IEEE Access, 2024.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Mmw-aqa: Multimodal in-the-wild dataset for action quality assess- ment.IEEE Access, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.724551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:24.438832Z digest=sha256:03c9f079ef4e4e7b05531a398496e20be46022434808eacd0c52cad73771ede0

Observation 5316e41a-8695-4090-a504-b1b094029013 · outbound

This paper cites Vision-language action knowledge learning for semantic-aware action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Vision-language action knowledge learning for semantic-aware action quality assessment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.608867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:24.554309Z digest=sha256:dbd0121de38de5637fd9bd59b09dc702180ceed883715346c80a6efc8497a242

Observation 373a5f44-f617-4a56-826b-dcf0f2b3c98a · outbound

This paper cites Learning semantics- guided representations for scoring figure skating.IEEE Transactions on Multimedia, 26:4987–4997, 2023.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning semantics- guided representations for scoring figure skating.IEEE Transactions on Multimedia, 26:4987–4997, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.453029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:24.665443Z digest=sha256:9ed8831f1640760ae3197a4f768fb2ee1d0f275ae70ae398ecc0b25b77dc5745

Observation 6466f5bb-3bea-4602-976b-3fd74bde4804 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal and cross-modal attention for audio-visual zero-shot learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.348648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:24.812348Z digest=sha256:d00cffef2a33b39556e0582331b73f205ac73541f85606138fad8006ab6acedc

Observation 28f464e8-07a7-4ac6-af2c-8f45d0800f4e · outbound

This paper cites Cross-attention is not always needed: Dynamic cross-attention for audio-visual dimensional emotion recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-attention is not always needed: Dynamic cross-attention for audio-visual dimensional emotion recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.201321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:24.981888Z digest=sha256:f6800a6be2583b8a6581ab3e228892fc1370bc019e1f3b76a50d52b231780283

Observation dc635a20-9a1d-41c8-b400-95e65f7d9f20 · outbound

This paper cites AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:18:26.702155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:25.094671Z digest=sha256:c39a96271eafe05515bc0cac18df27f66476a4afc8ce7fad80febd441acd43e1

Observation 81151f82-1911-418a-9873-2a844699de31 · outbound

This paper cites Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:25.205516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:25.205516Z digest=sha256:69d2a66318462b214e444c1d34792ac7c5978c6fd8e95e46963e93950a9aa388

Observation 9a091c2c-f28a-4b37-a51e-d7c82d18db2c · outbound

This paper cites Temporal alignment networks for long-term video.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal alignment networks for long-term video

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.997529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:25.294446Z digest=sha256:21134c1dbfb8bfc413b2ccaca06a66fdceb4b47a3aec0bd0fabe2e32be2193a5

Observation 0cce93b2-4bc4-4ec1-b98c-a11a30d0c358 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning.Advances in neural information processing systems, 35:38032–38045, 2022.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal and cross-modal attention for audio-visual zero-shot learning.Advances in neural information processing systems, 35:38032–38045, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.715831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:25.400298Z digest=sha256:ace714d4bb5f201d95d6407f6c5d0c17f6b3988fb392b37d19db8aab9e490a78

Observation 371a4223-90b5-496a-ad60-888a71e4af2e · outbound

This paper cites Video and accelerometer-based motion analysis for automated surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Video and accelerometer-based motion analysis for automated surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.491535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:25.469481Z digest=sha256:9331322bf909b2c7d81d72407b604f2ded87367be73eaa03cee46346e1930298

Observation be99137b-cb2b-4394-9f9e-23c44703000e · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audio set: An ontology and human-labeled dataset for audio events

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.040623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:25.735323Z digest=sha256:3433f0a568a1a5a403961a6a1e9996dab09e30284a1fd565e3dbaf202de548dc

Observation cae90976-fcc6-48fb-9880-545d7827f2bc · outbound

This paper cites Action quality assessment with temporal parsing transformer.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Action quality assessment with temporal parsing transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.837087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:25.826751Z digest=sha256:193e4cd22ad20189dd3533fc7b4543dd2b9fb72f346382d184491fa0e029be69

Observation 47c93c07-361c-486a-bf55-ee1634b7f01a · outbound

This paper cites Learning spatiotemporal features with 3d convolu- tional networks.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning spatiotemporal features with 3d convolu- tional networks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.690611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:25.887329Z digest=sha256:76321ff090e60fa6d8ac87356216ebec49fdf13b5117395b0d81578093546cd7

Observation a7d2812a-e2a4-4ec1-bf60-f120a2df2321 · outbound

This paper cites Video and accelerometer-based motion analysis for automated 43 surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Video and accelerometer-based motion analysis for automated 43 surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.509449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:25.948741Z digest=sha256:e2a41fb1bd97407279845ac434ed920b8ba501ee44207700a80f930aeb0dd39d

Observation be607b59-d4b6-4031-b12e-9f75cf801bc5 · outbound

This paper cites Deep resid- uallearningforimagerecognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Deep resid- uallearningforimagerecognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.296707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:26.008721Z digest=sha256:dabecfcc0e0624e31b673d4e7c6f44d2214ba2d4f58ceeaf9595f2128541551c

Observation b63b5445-ccb8-4bc5-b274-386495198d89 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Quo vadis, action recognition? a new model and the kinetics dataset

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.289591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:26.071958Z digest=sha256:009e3a489c5a4d96f53b52fdd23143722597159fba46b1041f0245726e1e2218

Observation 187f7f20-0af3-4e6f-a632-27b512db3adc · outbound

This paper cites AST: Audio Spectrogram Transformer.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment AST: Audio Spectrogram Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:26.116606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:26.116606Z digest=sha256:56187f19b0c27138c8fffb4d5c342cdd0c9c5aa878af46c55ca40e305b14fca3

Observation 7f0a990d-0458-4ace-99a2-e20b3a350c74 · outbound

This paper cites Joint visual and audio learning for video highlight detection.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Joint visual and audio learning for video highlight detection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.086584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:26.176463Z digest=sha256:81884992fdf2087e89af62133806a551924d5f76701d4246a20419c49f40affe

Observation da49cd3c-3a8c-444d-a127-ce5d3ba5b860 · outbound

This paper cites MSAF: Multimodal Split Attention Fusion.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment MSAF: Multimodal Split Attention Fusion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:26.255413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:26.255413Z digest=sha256:7dd5262834e7702e253dbc69b9e5215c479a30e7b2348e407d2b74e1e1fdc2a0

Observation be1ce44f-b478-4363-9b88-7c5c7f62d816 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:26.885434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:26.355829Z digest=sha256:bfe16bf64a8f1ecec55dfc7c134d3737b7c30173f6d5b937cfdceebd44fc840d

Pith citing papers

No inbound Pith citation observations are available.