Pith. sign in

Paper Citation Record · LEDGER

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA

As of 16 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:1908.03744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.03744 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:07:22.218205Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation edaa83b7-e2e2-414b-ba1f-b8f6525cfa04 · outbound

This paper cites YouTube-8M: A Large-Scale Video Classification Benchmark.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA YouTube-8M: A Large-Scale Video Classification Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.131966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.131966Z digest=sha256:5394d45803399ddcb7ad3e34773f787347f6802c5dbbfbb77285902cbbc24cd2

Observation 7e8daf12-f2fc-4bc7-b47b-277d4a194b9d · outbound

This paper cites Understanding affective content of music videos through learned representations.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Understanding affective content of music videos through learned representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.483212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.136248Z digest=sha256:f6946845d6e331215103de4fd58fa0d1ea347389bba00e974a8e70e4a381a208

Observation 7dc09abb-bf61-4d7c-adb9-4f432adb04e0 · outbound

This paper cites Deep canonical correlation analysis.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Deep canonical correlation analysis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.473211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.139764Z digest=sha256:9ac08f8e007fd413733f1aa419476313c33f9f26e9ba092a8f4efd3a6f6d67f9

Observation 7837889f-2622-48f0-a313-dcb664155dc3 · outbound

This paper cites The sound of an album cover: Probabilistic multimedia and information retrieval.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA The sound of an album cover: Probabilistic multimedia and information retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.463505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.143062Z digest=sha256:062642f39decfb1e04aa5993da2c1701d0bdcac85353cbe33830e9ebe479ddd0

Observation f7e0b30e-736d-40de-9116-be99f523e36e · outbound

This paper cites An introduction to support vector machines and other kernel-based learning methods.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA An introduction to support vector machines and other kernel-based learning methods

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.453809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.146556Z digest=sha256:304873483bfdb15f0bd4b9fce43f1493d587f0d5f695c1785fa135b566771423

Observation 3dff33df-3ffc-47bb-b821-310b50a8c8e2 · outbound

This paper cites Cross-modal retrieval with correspondence autoencoder.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Cross-modal retrieval with correspondence autoencoder

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.442823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.150546Z digest=sha256:4bb26e5f40ec3249796deeb37dec6fa812848605811d96cdd897276d78cd9a61

Observation 72938573-be2f-4005-9b99-f253a5dd90bc · outbound

This paper cites On the correlation of automatic audio and visual segmentations of music videos.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA On the correlation of automatic audio and visual segmentations of music videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.433267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.154219Z digest=sha256:97132facf206631f7aabfccd5a86715c8727a601a338bfd90db986e3bac606ce

Observation 8f310d66-a749-41d8-887d-7aad42f42bb8 · outbound

This paper cites Cnn architectures for large-scale audio classification.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Cnn architectures for large-scale audio classification

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.423368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.157331Z digest=sha256:302c12134e7cf972755c3aa1ad560b15f01a28ab96a07e07507757c2186126c8

Observation e1d37e28-0f59-49fd-8292-fc5f285a2c26 · outbound

This paper cites Long short-term memory.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Long short-term memory

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.160343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.160343Z digest=sha256:7a5efa7562e4e3e4b8b18380516eb00d7f17d4375c5717dc36bd014102a65f4d

Observation 028bf625-c259-4770-a099-ad9da76e79b3 · outbound

This paper cites Music thumbnail- ing via neural attention modeling of music emotion.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Music thumbnail- ing via neural attention modeling of music emotion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.407869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.163376Z digest=sha256:9de9d26699e4f70baed805c4b4c97cdb2010d8270e70deeaac2917a7cf6661f4

Observation f7eda8c6-2d04-478c-81ac-67c06bfd4f88 · outbound

This paper cites Deep cross-modal hashing.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Deep cross-modal hashing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.397516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.166491Z digest=sha256:dab35154a2248cf462bad9569ae920589562c6b43467a93c84d789dcaae57d25

Observation f147e022-b61a-4b38-a647-00bc8edafdf4 · outbound

This paper cites An Empirical Evaluation of doc2vec with Practical Insights into Document Embedding Generation.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA An Empirical Evaluation of doc2vec with Practical Insights into Document Embedding Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.169465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.169465Z digest=sha256:325bb3fd422c8a31eaccf48fb1d7efc396599500e0af755974d85b81399436aa

Observation 2f7c2056-da49-44cb-88d3-9d2cb30d71d6 · outbound

This paper cites Microsoft coco: Common objects in context.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Microsoft coco: Common objects in context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.173025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.173025Z digest=sha256:3ba65a62dd7dfa9450e992d965db93f80a49db056e787266bc5ecac7177eb9f8

Observation b14c390c-240a-47ba-94db-8099a66e412c · outbound

This paper cites Analysing the similarity of album art with self-organising maps.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Analysing the similarity of album art with self-organising maps

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.382359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.176259Z digest=sha256:2dd0011a384035074f5933f0d4065540ed400fd8c9aee4470032d06e3d8a11ec

Observation d1f5a63a-89a7-4658-8bc9-25906bee9b76 · outbound

This paper cites Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.179410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.179410Z digest=sha256:4f2b031a16649a719d65dcb9f551a1e5cd0581210b4a903fd46d4083656c6061

Observation 6f8a3219-647d-4497-a449-cb41a4f12a53 · outbound

This paper cites A new approach to cross-modal multimedia retrieval.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA A new approach to cross-modal multimedia retrieval

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.367457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.182353Z digest=sha256:2bbae69536a7656ae4cd074c7201a76a416f80d21f742724e4345bc1f8d0d2f8

Observation ebf0375b-db9e-4deb-bcf0-72f45119dc71 · outbound

This paper cites Cluster canonical correlation analysis.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Cluster canonical correlation analysis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.358840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.185564Z digest=sha256:91c5c0c4788cbba08d22174351e73eaf72ad89ac8ecafccd5e2eb9cee90f8c54

Observation 072cdb07-74bd-4a30-990c-9de01c336069 · outbound

This paper cites Advisor: Person- alized video soundtrack recommendation by late fusion with heuristic rankings.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Advisor: Person- alized video soundtrack recommendation by late fusion with heuristic rankings

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.350148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.188650Z digest=sha256:a7e8fd3f303e4c9ebfe18f57521a54b5905c17de98ae0bb608e8927ee9c78353

Observation 7a9cfa24-dc09-4e05-8cd2-98a0f9a5a1f7 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.192128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.192128Z digest=sha256:7660e048282bed541952651705178211e464bcac9f4b97964aa840dc0cbf2cf9

Observation 1a662c09-06b0-4fce-acc9-8baa0052e029 · outbound

This paper cites Canonical correlation analysis.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Canonical correlation analysis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.340817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.195799Z digest=sha256:5403171931d7e940cea348fe1d64e3292b8c8497834fd07932b5ad5c798589d7

Observation b98ac8e9-e402-45a3-9ae4-a1c9e39a7d5e · outbound

This paper cites Learning deep structure- preserving image-text embeddings.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Learning deep structure- preserving image-text embeddings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.331319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.198917Z digest=sha256:e673b406b2a0bce52ca0c85bbff0305eff36f99dea3881d8cf2915f5302c98e7

Observation a924905b-ff45-4798-9956-4b15d66ed2d0 · outbound

This paper cites Automatic music sound- track generation for outdoor videos from contextual sensor information.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Automatic music sound- track generation for outdoor videos from contextual sensor information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.321081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.202415Z digest=sha256:e74741cf28488976bc872b26f442a2dcb628b4905320eafd293d0e11690805aa

Observation 5db56d1a-31c2-4608-aced-c673483c5cca · outbound

This paper cites Venuenet: Fine-grained venue discovery by deep correlation learning.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Venuenet: Fine-grained venue discovery by deep correlation learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.311101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.205513Z digest=sha256:c1d361dc9d21cb12b9ba3ec726785904f6ad915b4f712de7e4fa62762034c240

Observation 34c9985f-f21a-44e4-b392-f4764cd7da2e · outbound

This paper cites Category- based deep cca for fine-grained venue discovery from multimodal data.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Category- based deep cca for fine-grained venue discovery from multimodal data

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.300623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.208647Z digest=sha256:92198f46348969495bf7e8107f71ef6c48f98edc6697d86017690d5efa50b00a

Observation 90f1c88c-60d5-4801-9d5e-8e9117e7d895 · outbound

This paper cites Deep cross- modal correlation learning for audio and lyrics in music retrieval.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Deep cross- modal correlation learning for audio and lyrics in music retrieval

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.290529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.211745Z digest=sha256:46839d10ddcd36ab9046da1caae81c3000f1f589a213b4a772084e1cf37f82ef

Observation f1aaf072-eba7-4ebe-a8fa-857543c9aab4 · outbound

This paper cites End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T14:07:22.214925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:07:22.214925Z digest=sha256:fb9a2b026eec357e8e2c3c7477196d7cfb6d8996d8a6c9e4ecdc55fccf538546

Observation 094eae7b-7d81-4126-a158-60d105936b4c · outbound

This paper cites Multi-view learning overview: Recent progress and new challenges.

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA Multi-view learning overview: Recent progress and new challenges

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:07:22.280948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:07:22.218205Z digest=sha256:85665d614b1d6bcbef9881c99436be20d0345f7ab0721370a2ecd9e4f54e029e

Pith citing papers

No inbound Pith citation observations are available.