Pith. sign in

Paper Citation Record · LEDGER

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.23524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23524 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:48.549041Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:45.412605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:48:48.744686Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be7dd9f4-c7b6-44a7-9938-a620c55f2b37 · outbound

This paper cites Tradi- tional methods primarily focus on identifying the most rele- vant videos based on a given query [1].

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Tradi- tional methods primarily focus on identifying the most rele- vant videos based on a given query [1]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:55.060702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:45.288073Z digest=sha256:c8b33ad074d994e836e138c9fdd19b27b58147032a25d9efabbfc06b69d6bcef

Observation b8a62b15-6959-49e3-a1fd-6abd14e3d672 · outbound

This paper cites CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:48.858700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:45.412605Z digest=sha256:6512083334b6f7269aa4d0fdf7b232877cd2a7281c0acf42ea1edcbe71bb7ca2

Observation b4972a19-a7db-48a8-b0dc-916c2c9ecc51 · outbound

This paper cites Pre-trained audio, VLP, and CBP feature extractors first extractF audio,F CBP , andF V LP features.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Pre-trained audio, VLP, and CBP feature extractors first extractF audio,F CBP , andF V LP features

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:54.769020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:45.570821Z digest=sha256:84c61e29152d7d835d97df65cbeaeef8ae8a84c4405faacd01398fd66abef881

Observation 7fc4dc23-0ae0-46ec-8056-d21829e2a281 · outbound

This paper cites Datasets THUMOS14consists of 200 validation and 213 test videos across 20 action classes, averaging 15 action segments per video.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Datasets THUMOS14consists of 200 validation and 213 test videos across 20 action classes, averaging 15 action segments per video

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:48:54.499112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:45.706456Z digest=sha256:6e16f1a8f33b61e666baa7d8d56243618afe635dfe763e7cedaac299a3c61382

Observation 9d21f908-8417-4162-9ee9-31e60b9c2cb9 · outbound

This paper cites This is the first time CLIP and audio are incorporated into UTAL.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization This is the first time CLIP and audio are incorporated into UTAL

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:54.205295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:45.858643Z digest=sha256:1df0a7f1d04939d127164b121d8192e2acd027a139c3ed782fbdf9e2802f690b

Observation 972e97a1-c7f6-466f-8b0d-dd926d4e3e0a · outbound

This paper cites an unresolved cited work.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:53.972067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:45.938182Z digest=sha256:9367c1db9411a4bdaaba68910316876bf8dd07b91f3f5da249fa8941006b8cb1

Observation 4ab952d0-c16a-479a-86a9-f0da1962b537 · outbound

This paper cites Two-stream consensus network for weakly-supervised temporal action local- ization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Two-stream consensus network for weakly-supervised temporal action local- ization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.658122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:46.145963Z digest=sha256:779aca0bdb6f0203355691c3caf3b782e888160d207aaf45592471d2262eb93c

Observation 0d566cc5-1f04-4187-a95c-7326fd47617b · outbound

This paper cites Weakly-supervised temporal action lo- calization by inferring salient snippet-feature,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly-supervised temporal action lo- calization by inferring salient snippet-feature,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.362935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:46.302868Z digest=sha256:a8120dbf43a5f863abb78a958f71499d6af54ad76491dcac8bbeb21cf2bbc6fa

Observation 9f980152-557e-422b-acdd-578e05553ac2 · outbound

This paper cites Two-stream networks for weakly-supervised temporal action local- ization with semantic-aware mechanisms,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Two-stream networks for weakly-supervised temporal action local- ization with semantic-aware mechanisms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.071181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:46.477326Z digest=sha256:c6e4891d84a2a41f04a13dedddba20a7d97c4648736e01b68151cfff16325bfe

Observation bedf9601-f849-45c1-893d-4433e4ec5471 · outbound

This paper cites Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.650668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:46.650668Z digest=sha256:6de5893cf6f9a658fded4cfb8b0eb65dc9ac614de3f8990188bc2eb96f3e0ad1

Observation 526873f1-ab5e-4d40-ad72-37b7d1b01cb6 · outbound

This paper cites Rethinking pseudo-label guided learning for weakly supervised temporal action localization from the perspective of noise correction,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Rethinking pseudo-label guided learning for weakly supervised temporal action localization from the perspective of noise correction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.740062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:46.819500Z digest=sha256:fa8825aafb9142f26138b73b57b92b725d6a02ba16652a1f783cecb0affe5f28

Observation 1eb1bdc3-2c6b-4836-badc-0f02d60d79f2 · outbound

This paper cites Apsl: Action-positive separation learning for unsupervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Apsl: Action-positive separation learning for unsupervised temporal action localization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.443787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:46.937683Z digest=sha256:1c31c07b0c2633145bf5c22390828c6755a356de57178225e77da7a9de5f601c

Observation e54e0186-25fe-479a-acec-d8f5c2c77a5d · outbound

This paper cites Learning temporal co-attention models for unsupervised video action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Learning temporal co-attention models for unsupervised video action localization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.029311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:47.187380Z digest=sha256:e29444e2c5af7b0e2f3e5960abe3741ab6d5474345769f217c5d58da78661aa7

Observation d4322ed8-362d-4821-a264-28948253a0cc · outbound

This paper cites Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.717507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:47.327506Z digest=sha256:6198a6b2848749d9efb3699aad314af4fd0a731e4fcb0d1adf3aeeb3b4a9f50f

Observation 1203afcb-20bb-4cbc-ae52-0b04790e2c0b · outbound

This paper cites Revisiting foreground and background separation in weakly-supervised temporal action local- ization: A clustering-based approach,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Revisiting foreground and background separation in weakly-supervised temporal action local- ization: A clustering-based approach,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.345450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:47.485299Z digest=sha256:e7c9d99d15f4dae288ab711200c56431f47a803a0d8889dc974cf23eea6fd1f6

Observation 2b287037-7dee-41e7-b1c1-deed347d2a52 · outbound

This paper cites Weakly supervised action localization by sparse temporal pooling network,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly supervised action localization by sparse temporal pooling network,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.957971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:47.600169Z digest=sha256:b7ee9e42c3dfb61133b0803272bae70ba4b6f1f80945b6dbec68e174e2f15b36

Observation 571033fb-cc77-4cb0-ae8d-758dfd04c236 · outbound

This paper cites Weakly-supervised action localization with background modeling,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly-supervised action localization with background modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.754831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:47.796538Z digest=sha256:8fdd988782f87e2cd2e2cdfe057643012f554b1fa71218e00681334fa1a834b9

Observation 3c64e08f-d674-44af-a81b-f7b95e026678 · outbound

This paper cites Proposal-based multiple instance learning for weakly-supervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Proposal-based multiple instance learning for weakly-supervised temporal action localization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.527963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:47.946877Z digest=sha256:e6f20b35a07a2c8f7865d0ef253bc43857aad1f6113a0ef86d7dfc8f073c7106

Observation 3d868298-5ba4-48a8-aeff-daaad0a00920 · outbound

This paper cites IMDPrompter: Adapting SAM to image manipulation detection by cross-view au- tomated prompt learning,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization IMDPrompter: Adapting SAM to image manipulation detection by cross-view au- tomated prompt learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.257347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:48.045654Z digest=sha256:83f3e406300a4ba6e4ee740947d78910591808513a03fbf629571f6b1fa8c908

Observation 1ce0dbff-495f-464a-9a6d-e5fbfb64174c · outbound

This paper cites Gim: A million-scale benchmark for generative image manipulation detection and localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Gim: A million-scale benchmark for generative image manipulation detection and localization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.843937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:48.218893Z digest=sha256:98fa27aa9b17f7e2c54e2ee50ab8d25725d0201bf5c4c318568222d6ef51b038

Observation af09fa38-d8d3-47cb-8882-7144e755d5c9 · outbound

This paper cites Seeing beyond noise: Joint graph structure evaluation and denoising for mul- timodal recommendation,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Seeing beyond noise: Joint graph structure evaluation and denoising for mul- timodal recommendation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.414467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:48.390480Z digest=sha256:3d0c80626ef4e56b73a585d8f153cee437dad7751ab6e4e1b8caa75e353e2847

Observation e236e4ff-3518-4f22-8806-362f13497fa1 · outbound

This paper cites Distilling semantic priors from sam to effi- cient image restoration models,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Distilling semantic priors from sam to effi- cient image restoration models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.123969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:48.549041Z digest=sha256:d626bfadd1f2b7dd4ad73a72b6b92f141318b76ee007001c9fa6da8ee8054604

Pith citing papers

Observation b8a62b15-6959-49e3-a1fd-6abd14e3d672 · inbound

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization cites this paper.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:48.858700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:48:45.412605Z digest=sha256:6512083334b6f7269aa4d0fdf7b232877cd2a7281c0acf42ea1edcbe71bb7ca2