Pith. sign in

Paper Citation Record · LEDGER

Learning Joint Embedding for Cross-Modal Retrieval

As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:1908.07673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.07673 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:16:24.593576Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d238aa03-55ca-4e6e-8129-71be889206dc · outbound

This paper cites Automatic music soundtrack generation for outdoor videos from contextual sensor information,.

Learning Joint Embedding for Cross-Modal Retrieval Automatic music soundtrack generation for outdoor videos from contextual sensor information,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.787979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.546392Z digest=sha256:35197b2148c573e9f894dbef2a14eda5418d234fb629b64dc39aa991c1c24f21

Observation 9b1efe76-7671-4a67-86df-9f1ab3ba81af · outbound

This paper cites Learning from between- class examples for deep sound recognition,.

Learning Joint Embedding for Cross-Modal Retrieval Learning from between- class examples for deep sound recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.773698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.551304Z digest=sha256:fb64d929bfb9776b99aefc1bdf37f5022b6f72bfd3f3e2b9068ce18df0356dd8

Observation e03d3be8-3935-4476-ac10-8b7ef7242f62 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Learning Joint Embedding for Cross-Modal Retrieval Audio set: An ontology and human-labeled dataset for audio events,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.760890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.555184Z digest=sha256:331693083571d9bcf7c846c626865371d3ba0464bcf858ee60aec7c8daf435c7

Observation b5fc887b-d16d-4dfb-a876-6bf09d20996e · outbound

This paper cites Learnable pooling with context gating for video classification,.

Learning Joint Embedding for Cross-Modal Retrieval Learnable pooling with context gating for video classification,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.748867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.559381Z digest=sha256:518fa190d3cd85815e88831129011ac00829d9a4d120b1c45f70966f6c87ff8b

Observation fc98061d-f465-4767-b891-f5e8b9ce0585 · outbound

This paper cites Canonical corre- lation analysis: An overview with application to learning methods,.

Learning Joint Embedding for Cross-Modal Retrieval Canonical corre- lation analysis: An overview with application to learning methods,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.736801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.563378Z digest=sha256:d355e8b4ec6fb69b134b9ce970c4479892b189c03f47a4ddbc6203c1f273963e

Observation 76428757-470b-479c-b050-2776bb5f2875 · outbound

This paper cites Deep cross-modal correlation learning for audio and lyrics in music retrieval,.

Learning Joint Embedding for Cross-Modal Retrieval Deep cross-modal correlation learning for audio and lyrics in music retrieval,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.723292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.567716Z digest=sha256:16f064330ded3d77b73baebe2af9225456ec1b05e03fbbd75d6b6b026bd2e4a2

Observation 5cedddd4-a1fc-4aff-8323-da6f08938d47 · outbound

This paper cites Deep canonical cor- relation analysis,.

Learning Joint Embedding for Cross-Modal Retrieval Deep canonical cor- relation analysis,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.709204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.571545Z digest=sha256:1c118626c0c248ea6ccf36b9776fc278b121a6bf35e2b68c73b50cbec57d1c8a

Observation ed8844c9-040a-4842-b7c6-67c28acdf6d1 · outbound

This paper cites Cluster canonical correlation analysis,.

Learning Joint Embedding for Cross-Modal Retrieval Cluster canonical correlation analysis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.694975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.575675Z digest=sha256:cd10f91beebab426309f0c2424884d641af086ae050d98ca0a0aaa553eafb87b

Observation c76dbd56-44b0-4b03-a822-71db1087e6af · outbound

This paper cites Category-based deep cca for fine-grained venue discovery from multimodal data,.

Learning Joint Embedding for Cross-Modal Retrieval Category-based deep cca for fine-grained venue discovery from multimodal data,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.680522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.580262Z digest=sha256:b6319ecc6d2922d9494a97b8234dbe221e4a90f5bc8c8d6dca1db89b457f5c78

Observation 4e9ec701-28e4-4067-bac8-fcba0c4e53a2 · outbound

This paper cites Visual to sound: Generating natural sound for videos in the wild,.

Learning Joint Embedding for Cross-Modal Retrieval Visual to sound: Generating natural sound for videos in the wild,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.666533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.584317Z digest=sha256:11f1e4f0e59ad7780e9c4b17bb0888f2465abfe037d184b7bf9a2a8359b50081

Observation 1e2aac7b-6bc8-415a-901f-15fff455c26d · outbound

This paper cites Audio-visual embedding for cross- modal music video retrieval through supervised deep cca.

Learning Joint Embedding for Cross-Modal Retrieval Audio-visual embedding for cross- modal music video retrieval through supervised deep cca

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:16:24.653526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.588888Z digest=sha256:71922188547c4344e7eb5b20bc78fe0e6a23a96d19609cfa7e4804a06a81505e

Observation eac51361-2138-49b5-9d5f-a0a244d53a8b · outbound

This paper cites Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-modal Retrieval.

Learning Joint Embedding for Cross-Modal Retrieval Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-modal Retrieval

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-14T12:16:24.638969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T12:16:24.593576Z digest=sha256:3664d3e68d9d4c7fe6fc18aa57596c3c3ec293de2803010ed14119705ae12ca3

Pith citing papers

No inbound Pith citation observations are available.