Pith. sign in

Paper Citation Record · LEDGER

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA

As of 16 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:1908.03744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.03744 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:07:22.218205Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation edaa83b7-e2e2-414b-ba1f-b8f6525cfa04 · outbound

This paper cites YouTube-8M: A Large-Scale Video Classification Benchmark.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA YouTube-8M: A Large-Scale Video Classification Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.131966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.131966Z digest=sha256:8d9637e08c9ccd78eeda71bb5fbf6787622fafd5650cdb4278d0b779b50b04e4

Observation 7e8daf12-f2fc-4bc7-b47b-277d4a194b9d · outbound

This paper cites Understanding affective content of music videos through learned representations.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Understanding affective content of music videos through learned representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.483212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.136248Z digest=sha256:f7ebe11fea0d411d98bfd9454bf7a68d9f9add6f7a49ed131a268c89c9908219

Observation 7dc09abb-bf61-4d7c-adb9-4f432adb04e0 · outbound

This paper cites Deep canonical correlation analysis.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Deep canonical correlation analysis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.473211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.139764Z digest=sha256:5fe949723b5f65860bc6f517ff01112c9e1327fe4705e55843f50f3a841d8489

Observation 7837889f-2622-48f0-a313-dcb664155dc3 · outbound

This paper cites The sound of an album cover: Probabilistic multimedia and information retrieval.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA The sound of an album cover: Probabilistic multimedia and information retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.463505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.143062Z digest=sha256:3efa8825b97d45987751116d14cecc6a1843be374085fb3a5aba435bc8341efc

Observation f7e0b30e-736d-40de-9116-be99f523e36e · outbound

This paper cites An introduction to support vector machines and other kernel-based learning methods.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA An introduction to support vector machines and other kernel-based learning methods

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.453809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.146556Z digest=sha256:4bace4d48f666aa3bd513b275c5c085724a20b86f8cb48dd3e696b932553d276

Observation 3dff33df-3ffc-47bb-b821-310b50a8c8e2 · outbound

This paper cites Cross-modal retrieval with correspondence autoencoder.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Cross-modal retrieval with correspondence autoencoder

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.442823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.150546Z digest=sha256:4b55d756b2257a5f0df559a05ef25ae0d80ac5c156b6921147f61366cd08f45f

Observation 72938573-be2f-4005-9b99-f253a5dd90bc · outbound

This paper cites On the correlation of automatic audio and visual segmentations of music videos.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA On the correlation of automatic audio and visual segmentations of music videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.433267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.154219Z digest=sha256:5d3f90f838b072caf0313d01b20b0f1d93db30dab4b7c80829e95356679210dc

Observation 8f310d66-a749-41d8-887d-7aad42f42bb8 · outbound

This paper cites Cnn architectures for large-scale audio classification.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Cnn architectures for large-scale audio classification

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.423368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.157331Z digest=sha256:70319101f1dd36a7da1562fdbc73b0223e9d28ca03346859d579e791b36d5685

Observation e1d37e28-0f59-49fd-8292-fc5f285a2c26 · outbound

This paper cites Long short-term memory.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Long short-term memory

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.160343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.160343Z digest=sha256:24767868650aa9ca31d1114c4655cfcdfd3868d57335f78b1d4e057030550b9e

Observation 028bf625-c259-4770-a099-ad9da76e79b3 · outbound

This paper cites Music thumbnail- ing via neural attention modeling of music emotion.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Music thumbnail- ing via neural attention modeling of music emotion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.407869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.163376Z digest=sha256:ef1f15a630627f92190e20e7d52bdb1613b3ff57a9333b0d2a66e591ae3ee1ed

Observation f7eda8c6-2d04-478c-81ac-67c06bfd4f88 · outbound

This paper cites Deep cross-modal hashing.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Deep cross-modal hashing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.397516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.166491Z digest=sha256:c24b123754d920cc4c25b0111faa339d375ab6ec73964bb84310152a8954fb0b

Observation f147e022-b61a-4b38-a647-00bc8edafdf4 · outbound

This paper cites An Empirical Evaluation of doc2vec with Practical Insights into Document Embedding Generation.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA An Empirical Evaluation of doc2vec with Practical Insights into Document Embedding Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.169465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.169465Z digest=sha256:2073a495bff5e56997a48c3f6e4fbb3e50cb11b30050c372ece0eeb1e167b377

Observation 2f7c2056-da49-44cb-88d3-9d2cb30d71d6 · outbound

This paper cites Microsoft coco: Common objects in context.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Microsoft coco: Common objects in context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.173025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.173025Z digest=sha256:6842041c4d44e2911f9ef1ec9b6909a1da9a4050850f0c017338533836d1e524

Observation b14c390c-240a-47ba-94db-8099a66e412c · outbound

This paper cites Analysing the similarity of album art with self-organising maps.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Analysing the similarity of album art with self-organising maps

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.382359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.176259Z digest=sha256:97166df84bdacd185cadfd8ff2da0b722fc15a3f554c510c518e45109b2e353e

Observation d1f5a63a-89a7-4658-8bc9-25906bee9b76 · outbound

This paper cites Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.179410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.179410Z digest=sha256:e69152b80374316401c3d85b0677cb097888ad9a66ea12b90f0728b0189c68fd

Observation 6f8a3219-647d-4497-a449-cb41a4f12a53 · outbound

This paper cites A new approach to cross-modal multimedia retrieval.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA A new approach to cross-modal multimedia retrieval

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.367457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.182353Z digest=sha256:c62c354bdd7a0310cfb98d8f6c52250650eb17c1fa4ff99b5f5b1a37dc145750

Observation ebf0375b-db9e-4deb-bcf0-72f45119dc71 · outbound

This paper cites Cluster canonical correlation analysis.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Cluster canonical correlation analysis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.358840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.185564Z digest=sha256:5c6c1ecde60ea3b93d7dc67c08a6830c1e451320e1467c0d4ec013e8149eea7a

Observation 072cdb07-74bd-4a30-990c-9de01c336069 · outbound

This paper cites Advisor: Person- alized video soundtrack recommendation by late fusion with heuristic rankings.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Advisor: Person- alized video soundtrack recommendation by late fusion with heuristic rankings

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.350148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.188650Z digest=sha256:21191113c66418c184b6acb6882e47d9ff1dbcc99c393dfc5dce0e9bf0343efd

Observation 7a9cfa24-dc09-4e05-8cd2-98a0f9a5a1f7 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.192128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.192128Z digest=sha256:33b7b30cfac2443aac84317d8d4e4283e5cba3193cb85aae905ef16f8236f4a8

Observation 1a662c09-06b0-4fce-acc9-8baa0052e029 · outbound

This paper cites Canonical correlation analysis.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Canonical correlation analysis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.340817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.195799Z digest=sha256:e76149773258602f0ea9e54816d64307ce58751ee3150d5f9c68fc4f8c34dd33

Observation b98ac8e9-e402-45a3-9ae4-a1c9e39a7d5e · outbound

This paper cites Learning deep structure- preserving image-text embeddings.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Learning deep structure- preserving image-text embeddings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.331319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.198917Z digest=sha256:36b59bbe761ff5c922aeb68aaf4b93ef3c18f7ac0ecdaf7e00d477986320de65

Observation a924905b-ff45-4798-9956-4b15d66ed2d0 · outbound

This paper cites Automatic music sound- track generation for outdoor videos from contextual sensor information.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Automatic music sound- track generation for outdoor videos from contextual sensor information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.321081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.202415Z digest=sha256:5c77cbcd996ac3feb60346373db4df934c44bff53bf3d59da1e48828fb734fe3

Observation 5db56d1a-31c2-4608-aced-c673483c5cca · outbound

This paper cites Venuenet: Fine-grained venue discovery by deep correlation learning.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Venuenet: Fine-grained venue discovery by deep correlation learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.311101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.205513Z digest=sha256:f9a7c36a8def3e5aa31482982546fe822095a7bb37b9f3cac95d7ae92e4c6d6e

Observation 34c9985f-f21a-44e4-b392-f4764cd7da2e · outbound

This paper cites Category- based deep cca for fine-grained venue discovery from multimodal data.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Category- based deep cca for fine-grained venue discovery from multimodal data

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.300623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.208647Z digest=sha256:56fb766f003254742d15806d7b65da63b0395f2c15b69bf55342cf62da6a52bd

Observation 90f1c88c-60d5-4801-9d5e-8e9117e7d895 · outbound

This paper cites Deep cross- modal correlation learning for audio and lyrics in music retrieval.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Deep cross- modal correlation learning for audio and lyrics in music retrieval

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.290529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.211745Z digest=sha256:eacc2cbb0ab4e130db9e2a620a2d0a0d74afc5503e04773cbdae0fff2758923e

Observation f1aaf072-eba7-4ebe-a8fa-857543c9aab4 · outbound

This paper cites End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.214925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.214925Z digest=sha256:9760d6c3df0535522fcb307c42ac31b72b5838277f7e3f88ecc7e05b2297ca39

Observation 094eae7b-7d81-4126-a158-60d105936b4c · outbound

This paper cites Multi-view learning overview: Recent progress and new challenges.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Multi-view learning overview: Recent progress and new challenges

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.280948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:07:22.218205Z digest=sha256:d37a09e056949d857d72180ed8f355b3f49e6afe2f313936e98a12ffbad43498

Pith citing papers

No inbound Pith citation observations are available.