Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:48.549041Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.23524.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:48.549041Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:45.412605Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:48:48.744686Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation be7dd9f4-c7b6-44a7-9938-a620c55f2b37 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Tradi- tional methods primarily focus on identifying the most rele- vant videos based on a given query [1]
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8a62b15-6959-49e3-a1fd-6abd14e3d672 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4972a19-a7db-48a8-b0dc-916c2c9ecc51 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Pre-trained audio, VLP, and CBP feature extractors first extractF audio,F CBP , andF V LP features
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7fc4dc23-0ae0-46ec-8056-d21829e2a281 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Datasets THUMOS14consists of 200 validation and 213 test videos across 20 action classes, averaging 15 action segments per video
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d21f908-8417-4162-9ee9-31e60b9c2cb9 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization This is the first time CLIP and audio are incorporated into UTAL
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 972e97a1-c7f6-466f-8b0d-dd926d4e3e0a · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ab952d0-c16a-479a-86a9-f0da1962b537 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Two-stream consensus network for weakly-supervised temporal action local- ization,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d566cc5-1f04-4187-a95c-7326fd47617b · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly-supervised temporal action lo- calization by inferring salient snippet-feature,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f980152-557e-422b-acdd-578e05553ac2 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Two-stream networks for weakly-supervised temporal action local- ization with semantic-aware mechanisms,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bedf9601-f849-45c1-893d-4433e4ec5471 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526873f1-ab5e-4d40-ad72-37b7d1b01cb6 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Rethinking pseudo-label guided learning for weakly supervised temporal action localization from the perspective of noise correction,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1eb1bdc3-2c6b-4836-badc-0f02d60d79f2 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Apsl: Action-positive separation learning for unsupervised temporal action localization,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e54e0186-25fe-479a-acec-d8f5c2c77a5d · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Learning temporal co-attention models for unsupervised video action localization,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4322ed8-362d-4821-a264-28948253a0cc · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1203afcb-20bb-4cbc-ae52-0b04790e2c0b · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Revisiting foreground and background separation in weakly-supervised temporal action local- ization: A clustering-based approach,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b287037-7dee-41e7-b1c1-deed347d2a52 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly supervised action localization by sparse temporal pooling network,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 571033fb-cc77-4cb0-ae8d-758dfd04c236 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly-supervised action localization with background modeling,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c64e08f-d674-44af-a81b-f7b95e026678 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Proposal-based multiple instance learning for weakly-supervised temporal action localization,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d868298-5ba4-48a8-aeff-daaad0a00920 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization IMDPrompter: Adapting SAM to image manipulation detection by cross-view au- tomated prompt learning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ce0dbff-495f-464a-9a6d-e5fbfb64174c · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Gim: A million-scale benchmark for generative image manipulation detection and localization,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af09fa38-d8d3-47cb-8882-7144e755d5c9 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Seeing beyond noise: Joint graph structure evaluation and denoising for mul- timodal recommendation,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e236e4ff-3518-4f22-8806-362f13497fa1 · outbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Distilling semantic priors from sam to effi- cient image restoration models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8a62b15-6959-49e3-a1fd-6abd14e3d672 · inbound
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.