Pith. sign in

Paper Citation Record · LEDGER

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.23524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23524 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:48.549041Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:45.412605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:48:48.744686Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be7dd9f4-c7b6-44a7-9938-a620c55f2b37 · outbound

This paper cites Tradi- tional methods primarily focus on identifying the most rele- vant videos based on a given query [1].

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Tradi- tional methods primarily focus on identifying the most rele- vant videos based on a given query [1]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:55.060702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:45.288073Z digest=sha256:52fcad1b93fea956a35af2ed97e2573c6cb3b2eaf6ae9f48f1b51b6e4bca5fa2

Observation b8a62b15-6959-49e3-a1fd-6abd14e3d672 · outbound

This paper cites CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:48.858700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:45.412605Z digest=sha256:017148e4fd5ff7d0d95bcbf36ebc0877fd36e981159c2b07422ca4f482701405

Observation b4972a19-a7db-48a8-b0dc-916c2c9ecc51 · outbound

This paper cites Pre-trained audio, VLP, and CBP feature extractors first extractF audio,F CBP , andF V LP features.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Pre-trained audio, VLP, and CBP feature extractors first extractF audio,F CBP , andF V LP features

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:54.769020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:45.570821Z digest=sha256:5e4b341c62c8fd4c231135c8ac246ab61f181339151775a75b2f26dfae4e3a19

Observation 7fc4dc23-0ae0-46ec-8056-d21829e2a281 · outbound

This paper cites Datasets THUMOS14consists of 200 validation and 213 test videos across 20 action classes, averaging 15 action segments per video.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Datasets THUMOS14consists of 200 validation and 213 test videos across 20 action classes, averaging 15 action segments per video

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:48:54.499112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:45.706456Z digest=sha256:bc45f658211496f034bd7017c42fb1f91b3c7b8bad86fd8dab6c0c0a04eeaff9

Observation 9d21f908-8417-4162-9ee9-31e60b9c2cb9 · outbound

This paper cites This is the first time CLIP and audio are incorporated into UTAL.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization This is the first time CLIP and audio are incorporated into UTAL

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:54.205295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:45.858643Z digest=sha256:5880407b6b297f25e6fb066dbc55264ce2774c135a2b499a91403eedf62b43d2

Observation 972e97a1-c7f6-466f-8b0d-dd926d4e3e0a · outbound

This paper cites an unresolved cited work.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:53.972067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:45.938182Z digest=sha256:4853dab57acd06566e50c82c5d0643e7f79353da6c50035e7b3e31f4ff0d1ffc

Observation 4ab952d0-c16a-479a-86a9-f0da1962b537 · outbound

This paper cites Two-stream consensus network for weakly-supervised temporal action local- ization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Two-stream consensus network for weakly-supervised temporal action local- ization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.658122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:46.145963Z digest=sha256:e65be37c2dc3ffe1c1bd5e57a202481f2bb00e339868912329233d00a87c7c46

Observation 0d566cc5-1f04-4187-a95c-7326fd47617b · outbound

This paper cites Weakly-supervised temporal action lo- calization by inferring salient snippet-feature,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly-supervised temporal action lo- calization by inferring salient snippet-feature,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.362935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:46.302868Z digest=sha256:9c88fed477b53e0a9435593778af7855ff857b825855ea9546406ab598fcc8bf

Observation 9f980152-557e-422b-acdd-578e05553ac2 · outbound

This paper cites Two-stream networks for weakly-supervised temporal action local- ization with semantic-aware mechanisms,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Two-stream networks for weakly-supervised temporal action local- ization with semantic-aware mechanisms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.071181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:46.477326Z digest=sha256:f14bc20c554972294e9244bb1be0d24da946084c1f847881390e0ef57164bddc

Observation bedf9601-f849-45c1-893d-4433e4ec5471 · outbound

This paper cites Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.650668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:46.650668Z digest=sha256:6de5893cf6f9a658fded4cfb8b0eb65dc9ac614de3f8990188bc2eb96f3e0ad1

Observation 526873f1-ab5e-4d40-ad72-37b7d1b01cb6 · outbound

This paper cites Rethinking pseudo-label guided learning for weakly supervised temporal action localization from the perspective of noise correction,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Rethinking pseudo-label guided learning for weakly supervised temporal action localization from the perspective of noise correction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.740062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:46.819500Z digest=sha256:b45ef958763b8f5e803fd1aa496cac279287b6b59c948d22282af3a42bb15811

Observation 1eb1bdc3-2c6b-4836-badc-0f02d60d79f2 · outbound

This paper cites Apsl: Action-positive separation learning for unsupervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Apsl: Action-positive separation learning for unsupervised temporal action localization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.443787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:46.937683Z digest=sha256:45caaef5fa187ce00d11e5be505dc8d965d0d90db36fb9a43a97bac88b8f918e

Observation e54e0186-25fe-479a-acec-d8f5c2c77a5d · outbound

This paper cites Learning temporal co-attention models for unsupervised video action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Learning temporal co-attention models for unsupervised video action localization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.029311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:47.187380Z digest=sha256:8ccc544454ba5cdb11590c29497a3b0649cd9bb8151a35b15b90b05631817294

Observation d4322ed8-362d-4821-a264-28948253a0cc · outbound

This paper cites Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.717507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:47.327506Z digest=sha256:4b82d1ce309f1e9343731e735063fc118a005dfd5be792d73488da0032d4306e

Observation 1203afcb-20bb-4cbc-ae52-0b04790e2c0b · outbound

This paper cites Revisiting foreground and background separation in weakly-supervised temporal action local- ization: A clustering-based approach,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Revisiting foreground and background separation in weakly-supervised temporal action local- ization: A clustering-based approach,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.345450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:47.485299Z digest=sha256:c367fd27c0a5ab5a50b833883321116f10684eddc7c9cd1060723cea8d1f2556

Observation 2b287037-7dee-41e7-b1c1-deed347d2a52 · outbound

This paper cites Weakly supervised action localization by sparse temporal pooling network,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly supervised action localization by sparse temporal pooling network,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.957971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:47.600169Z digest=sha256:b5c890e793ee84978d8be9266e5171b96f3fc3750b1f34205be326aff9a3679c

Observation 571033fb-cc77-4cb0-ae8d-758dfd04c236 · outbound

This paper cites Weakly-supervised action localization with background modeling,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly-supervised action localization with background modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.754831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:47.796538Z digest=sha256:b5121cc3c54695e167141a8855347bb1c5b57539a2f63e7c6efc766fc46314a4

Observation 3c64e08f-d674-44af-a81b-f7b95e026678 · outbound

This paper cites Proposal-based multiple instance learning for weakly-supervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Proposal-based multiple instance learning for weakly-supervised temporal action localization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.527963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:47.946877Z digest=sha256:b8b58bcc08d4b68ccb85d1d7ac67918020ac07e8272977120d3010fd35a9918c

Observation 3d868298-5ba4-48a8-aeff-daaad0a00920 · outbound

This paper cites IMDPrompter: Adapting SAM to image manipulation detection by cross-view au- tomated prompt learning,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization IMDPrompter: Adapting SAM to image manipulation detection by cross-view au- tomated prompt learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.257347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:48.045654Z digest=sha256:f299643f9b167b56b8f81ad739a49f3d4c829fdfe30d939960a04c36751541b1

Observation 1ce0dbff-495f-464a-9a6d-e5fbfb64174c · outbound

This paper cites Gim: A million-scale benchmark for generative image manipulation detection and localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Gim: A million-scale benchmark for generative image manipulation detection and localization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.843937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:48.218893Z digest=sha256:a7dd850743ca9dd40282a27fc848e17283f96abcae77b41c04802fe171006c1f

Observation af09fa38-d8d3-47cb-8882-7144e755d5c9 · outbound

This paper cites Seeing beyond noise: Joint graph structure evaluation and denoising for mul- timodal recommendation,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Seeing beyond noise: Joint graph structure evaluation and denoising for mul- timodal recommendation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.414467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:48.390480Z digest=sha256:2af754ac2c74b7131fe6ee323620a947be7bd52c05acc2a831308756911c9a77

Observation e236e4ff-3518-4f22-8806-362f13497fa1 · outbound

This paper cites Distilling semantic priors from sam to effi- cient image restoration models,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Distilling semantic priors from sam to effi- cient image restoration models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.123969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:48.549041Z digest=sha256:b2311d93ed92aa2116314d7d24cc9153325af535a1b8217ad7120f1b46851c52

Pith citing papers

Observation b8a62b15-6959-49e3-a1fd-6abd14e3d672 · inbound

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization cites this paper.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:48.858700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:48:45.412605Z digest=sha256:017148e4fd5ff7d0d95bcbf36ebc0877fd36e981159c2b07422ca4f482701405