Pith. sign in

Paper Citation Record · LEDGER

Learning Joint Embedding for Cross-Modal Retrieval

As of 16 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:1908.07673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.07673 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:16:24.593576Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d238aa03-55ca-4e6e-8129-71be889206dc · outbound

This paper cites Automatic music soundtrack generation for outdoor videos from contextual sensor information,.

Learning Joint Embedding for Cross-Modal Retrieval Automatic music soundtrack generation for outdoor videos from contextual sensor information,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.787979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.546392Z digest=sha256:297d89fe2db5a540acd44923361edd72c708c8e7168e18ad0560c9425caba015

Observation 9b1efe76-7671-4a67-86df-9f1ab3ba81af · outbound

This paper cites Learning from between- class examples for deep sound recognition,.

Learning Joint Embedding for Cross-Modal Retrieval Learning from between- class examples for deep sound recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.773698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.551304Z digest=sha256:9c6aa310c162f088cb493509048c82058c40dc08910f8225924fc6acf4f2b28c

Observation e03d3be8-3935-4476-ac10-8b7ef7242f62 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Learning Joint Embedding for Cross-Modal Retrieval Audio set: An ontology and human-labeled dataset for audio events,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.760890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.555184Z digest=sha256:e678b4d9675278d2e2831edbbb1edf6b7d7386afc7d2ce9a301b645526a2ad57

Observation b5fc887b-d16d-4dfb-a876-6bf09d20996e · outbound

This paper cites Learnable pooling with context gating for video classification,.

Learning Joint Embedding for Cross-Modal Retrieval Learnable pooling with context gating for video classification,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.748867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.559381Z digest=sha256:9f575c641a6ff4228bd930a6db709147f7a0b54a6d230d2a996264c704847d0b

Observation fc98061d-f465-4767-b891-f5e8b9ce0585 · outbound

This paper cites Canonical corre- lation analysis: An overview with application to learning methods,.

Learning Joint Embedding for Cross-Modal Retrieval Canonical corre- lation analysis: An overview with application to learning methods,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.736801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.563378Z digest=sha256:a6c4fbc703a3b83df8fcc6ee1e0406bd3ad2bc548ca370cbed1b43a9cc431c15

Observation 76428757-470b-479c-b050-2776bb5f2875 · outbound

This paper cites Deep cross-modal correlation learning for audio and lyrics in music retrieval,.

Learning Joint Embedding for Cross-Modal Retrieval Deep cross-modal correlation learning for audio and lyrics in music retrieval,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.723292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.567716Z digest=sha256:f465dc4b9f520d0c4dea29c6ba32ffabe3e7753c6a6e61c44172266ae807a445

Observation 5cedddd4-a1fc-4aff-8323-da6f08938d47 · outbound

This paper cites Deep canonical cor- relation analysis,.

Learning Joint Embedding for Cross-Modal Retrieval Deep canonical cor- relation analysis,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.709204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.571545Z digest=sha256:1caea7cc5173893f670d0fdc74cce4eb131757ce5a43d2019a73f5605938e6b3

Observation ed8844c9-040a-4842-b7c6-67c28acdf6d1 · outbound

This paper cites Cluster canonical correlation analysis,.

Learning Joint Embedding for Cross-Modal Retrieval Cluster canonical correlation analysis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.694975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.575675Z digest=sha256:6c5986c1d6d02ad005a29f0ccb909c0af4336ecaae2d3e5b0c9df66e4cf67fcf

Observation c76dbd56-44b0-4b03-a822-71db1087e6af · outbound

This paper cites Category-based deep cca for fine-grained venue discovery from multimodal data,.

Learning Joint Embedding for Cross-Modal Retrieval Category-based deep cca for fine-grained venue discovery from multimodal data,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.680522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.580262Z digest=sha256:f239546731cd005923c9d4f8df96eb10c5ce55f2503e0e786e7bf9cca7ca9bba

Observation 4e9ec701-28e4-4067-bac8-fcba0c4e53a2 · outbound

This paper cites Visual to sound: Generating natural sound for videos in the wild,.

Learning Joint Embedding for Cross-Modal Retrieval Visual to sound: Generating natural sound for videos in the wild,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.666533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.584317Z digest=sha256:3551e5c2c26f63ef48afa3bfef63953e84a16d2d46422b200a607dc867511093

Observation 1e2aac7b-6bc8-415a-901f-15fff455c26d · outbound

This paper cites Audio-visual embedding for cross- modal music video retrieval through supervised deep cca.

Learning Joint Embedding for Cross-Modal Retrieval Audio-visual embedding for cross- modal music video retrieval through supervised deep cca

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.653526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.588888Z digest=sha256:da0c7558ffceea4547805350c7bb6c2fcd99cfe02eada79b92d00c7ccaa20ad0

Observation eac51361-2138-49b5-9d5f-a0a244d53a8b · outbound

This paper cites Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-modal Retrieval.

Learning Joint Embedding for Cross-Modal Retrieval Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-modal Retrieval

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-14T12:16:24.638969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:16:24.593576Z digest=sha256:3180a2e2b190bd77f7b10d3c8557b06b78e56b2ca308e43c2598c4e628241612

Pith citing papers

No inbound Pith citation observations are available.