Pith. sign in

Paper Citation Record · LEDGER

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

As of 15 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2505.17002.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17002 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:54:27.753449Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:54:25.077226Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:19:53.890089Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e2ae54f-a40b-4385-8d13-ed1353858144 · outbound

This paper cites This task is brought into the com- puter vision community by Nagrani et al.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association This task is brought into the com- puter vision community by Nagrani et al

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:32.745213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.037032Z digest=sha256:00934af1c40839a32c6ab61b3039da81e1da5bb9945007e5972684e6e2a05aa8

Observation 19de554e-2010-46a2-95f3-8195cab4a911 · outbound

This paper cites PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:54:25.077226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:54:25.077226Z digest=sha256:d198a34ceead503527872fe07ad4371ccc823ebb51891db1ad69eccaf00f7984

Observation 5c261a68-bb76-450d-8130-d6c5bc61429d · outbound

This paper cites Baseline Approach In this work, we adapt a two-branch framework as our base- line method to establish face-voice association [10].

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Baseline Approach In this work, we adapt a two-branch framework as our base- line method to establish face-voice association [10]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:32.620254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.132483Z digest=sha256:1e47d290fded019c765f70d7c7c6cda88e76c4540a582786d9e5ec74467097e9

Observation e793fa15-e583-447c-9324-f5f824f2889c · outbound

This paper cites We set the initial learning rate to 2e−5.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association We set the initial learning rate to 2e−5

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:32.494309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.222063Z digest=sha256:af782d14e2da10da87318aef0d0d306bec408764a8424e836c5a6c06fa24ed79

Observation 72f4a298-40dd-4618-9576-a26a70d3aca1 · outbound

This paper cites We demonstrated that the precise alignment of features is a crucial step for obtaining better performance in the face-voice associa- tion task.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association We demonstrated that the precise alignment of features is a crucial step for obtaining better performance in the face-voice associa- tion task

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:32.321517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.287207Z digest=sha256:93f04f2de86ba93c70378cbf886a8d20abbaa68a3dbcb42110400cebe1a19b6d

Observation ac889a9d-04df-4567-bd1f-ddef20f445d7 · outbound

This paper cites Putting the face to the voice’: Matching identity across modal- ity,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Putting the face to the voice’: Matching identity across modal- ity,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:32.203762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.352083Z digest=sha256:6ae25a348f4b7d69c80b3131b3e751ac78ce905cccc0d694f9f79a15db6f8e1b

Observation 233bfe72-400b-4ffc-bab6-85b60553e5c8 · outbound

This paper cites Thinking the voice: neu- ral correlates of voice perception,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Thinking the voice: neu- ral correlates of voice perception,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:32.107457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.412953Z digest=sha256:694e68865064079334fddf3b9cf8ec372d5f5d9b5669e7d2b79352e1b2545e30

Observation 6f673626-0f95-41d2-8f85-942fb463e984 · outbound

This paper cites Hearing a face: Cross-modal speaker matching using isolated visible speech,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Hearing a face: Cross-modal speaker matching using isolated visible speech,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:32.004254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.480176Z digest=sha256:2b8ce6aa9747b4db6890892f76d5e034d44f95433b0383df7fddc53170188a6a

Observation ec56badc-d388-4900-9932-f68574137239 · outbound

This paper cites Seeing voices and hearing faces: Cross-modal biometric matching,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Seeing voices and hearing faces: Cross-modal biometric matching,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:31.874448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.558164Z digest=sha256:6643da89d573bdecf7e9d45a6f44b1f8631b633b322da0bd786c4db91f82be5d

Observation ff51b706-4729-4a51-9570-56c018c7c215 · outbound

This paper cites Learnable pins: Cross-modal embeddings for person identity,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Learnable pins: Cross-modal embeddings for person identity,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:31.722720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.648218Z digest=sha256:9da82196f1a1cdaa3476eeeede2d6e93cc195676e5e7652c16471622c942cb95

Observation 72b7d7d3-5283-470d-b23a-49e7a3758935 · outbound

This paper cites Face-voice match- ing using cross-modal embeddings,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Face-voice match- ing using cross-modal embeddings,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:31.597629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.764855Z digest=sha256:a31d171fdc69e41eb3aa76f694f305dfcf406edbba67a2ebd281d0d25eb79626

Observation d4f0024d-0421-45e7-aba6-d40763571db6 · outbound

This paper cites Dis- joint mapping network for cross-modal matching of voices and faces,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Dis- joint mapping network for cross-modal matching of voices and faces,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:31.453554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.840801Z digest=sha256:1a1af017c2a0db465fddb103826851a7851e51dae19cf60c94311a59f7c54c3b

Observation ebab6516-4788-4cc0-9ff9-5f605be841bf · outbound

This paper cites Deep latent space learning for cross-modal mapping of audio and visual signals,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Deep latent space learning for cross-modal mapping of audio and visual signals,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:31.307401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:25.935750Z digest=sha256:37080fe0cec8e13de0026543600458324009be9c0d2be4ab26f9196ffab7c4a8

Observation 3c8af93b-5831-4b11-858a-b3d426a1dfe3 · outbound

This paper cites Seeking the shape of sound: An adaptive framework for learning voice- face association,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Seeking the shape of sound: An adaptive framework for learning voice- face association,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:31.166456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.025224Z digest=sha256:4577cf1717ee041ea6c578e2213b42647f47bf1087b02e0cb648f05505cbfba8

Observation a3d21471-1f8e-4080-b265-87d36b223760 · outbound

This paper cites Fusion and orthogonal projection for improved face-voice association,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Fusion and orthogonal projection for improved face-voice association,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:31.023319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.068768Z digest=sha256:97476f32ddaba6ee675cff0e3d07ee9958c512e9944c8031e66c5921bf1d17f4

Observation 52758eb0-4474-4369-a6a1-72beea0779e6 · outbound

This paper cites Single-branch network for multimodal training,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Single-branch network for multimodal training,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:30.910744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.157777Z digest=sha256:b71ec28bacb721f7eae727a19c77ad9516f5e9a9e12b67451b42dc2ea9567c69

Observation c3175f51-5922-4e1a-b149-4b48c2235a84 · outbound

This paper cites A synopsis of fame 2024 challenge: Associating faces with voices in multilingual environments,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association A synopsis of fame 2024 challenge: Associating faces with voices in multilingual environments,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:30.793467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.227767Z digest=sha256:cdca014880199421911d6d9b231a2e90de8a9f782d583310add8f911a06beeba

Observation bee28be6-c930-4fb7-ab16-e6a128b0ec48 · outbound

This paper cites Speaker recognition in realistic scenario using multimodal data,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Speaker recognition in realistic scenario using multimodal data,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:30.632867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.300480Z digest=sha256:dd6aa82fbdda9c6bd76e776b984680fb1cb51bd1939550bccc4d2ae7696b1583

Observation be741695-1fa3-484d-ab92-e8a30bc31676 · outbound

This paper cites Cross-modal speaker verification and recognition: A multilingual perspective,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Cross-modal speaker verification and recognition: A multilingual perspective,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:30.482701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.358285Z digest=sha256:7aef5fa9247011c555c8ec266a04765f5b5502abab22fffb883e9a51b2a6822d

Observation cb6b7d32-c97b-4194-9bf6-791f98ea4829 · outbound

This paper cites Local-global contrast for learning voice-face representations,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Local-global contrast for learning voice-face representations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:30.303470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.422008Z digest=sha256:fd925329ef6ae1c64406490b6268d3ad8a7545c38fb65b8ab346b65f751758ba

Observation 1cf49cc9-ab3c-4a68-90f8-058a2ae7cbec · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association VoxCeleb: a large-scale speaker identification dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:54:26.487729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:54:26.487729Z digest=sha256:241f33a5b189d00b82b9a1a72c4ff3edef75c58561dc6740873a2ef91b835d4b

Observation 2b34aa0b-a7fb-48ad-b253-ea3c6c0b0b1c · outbound

This paper cites Multimodal ma- chine learning: A survey and taxonomy,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Multimodal ma- chine learning: A survey and taxonomy,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:30.135954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.570882Z digest=sha256:e8663d0b2d712b1dcd189421189e3c3c781dfd1abdfcc2db46269f7b3ed25c4a

Observation 589af4a3-9fdb-44e6-b4eb-68879e93630f · outbound

This paper cites Multimodal intelligence: Representation learning, information fusion, and applications,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Multimodal intelligence: Representation learning, information fusion, and applications,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:29.959366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.641126Z digest=sha256:5051b76bec4d789050f3181ae64396c7fabc92173144925ae86d2ee0165a9066

Observation c14a2ae2-b950-441e-90be-eb4de3dcdc5c · outbound

This paper cites Multimodal learning with trans- formers: A survey,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Multimodal learning with trans- formers: A survey,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:29.832815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.729811Z digest=sha256:3a8b664dcd4e269c989cd67149a6ef8c4d6f646b7036668c8e3caa0fd2f98c0a

Observation 68f3fafa-8144-4449-80fb-db3481cf8276 · outbound

This paper cites Guiding at- tention using partial-order relationships for image captioning,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Guiding at- tention using partial-order relationships for image captioning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:29.691377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.797790Z digest=sha256:6f3c979b040106225d90e701edf667f4c66d1a149e98707bf3fed584f84dde3b

Observation 47864025-a6ca-4e75-9956-f2b81e1292dc · outbound

This paper cites Dctm: Dilated convolutional trans- former model for multimodal engagement estimation in conversa- tion,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Dctm: Dilated convolutional trans- former model for multimodal engagement estimation in conversa- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:29.518762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.868920Z digest=sha256:8ba588dd7dc55666a9d5f471efa7fa8b99feef918be21d34c134b5c43c92d642

Observation a85c331a-8e54-4a6b-94e7-3ada0045d66c · outbound

This paper cites Representation tradeoffs for hyperbolic embeddings,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Representation tradeoffs for hyperbolic embeddings,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:29.355771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.921999Z digest=sha256:d6fb9f8ff356a7ed1f2dc18d8b719a9bd3e34d6ed7b7a3d0a62dbbba3d1325c9

Observation efb3c120-12e7-442e-8279-7b6e2066fa5a · outbound

This paper cites Hyperbolic image embeddings,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Hyperbolic image embeddings,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:29.194944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:26.996811Z digest=sha256:f3a2203ca48e4e77ba8265fccb26415382036ed661c5c8a8d2d0acc81f66089e

Observation cc9f730a-504d-4322-8f00-5fa3399829ce · outbound

This paper cites Accept the modality gap: An exploration in the hyperbolic space,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Accept the modality gap: An exploration in the hyperbolic space,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:29.039703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:27.070214Z digest=sha256:ddc6db3a233199c2d1d95a9efc95a4d1867f86b5d48a65c4add2142b5534a913

Observation d829e900-f005-4403-bc6f-61b185f9c551 · outbound

This paper cites Cross-modal scalable hyperbolic hi- erarchical clustering,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Cross-modal scalable hyperbolic hi- erarchical clustering,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:28.861382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:27.140338Z digest=sha256:7b9e18f06b49ca4512b49716a05ac6daa059108ec41a50cb346f1701e1782b22

Observation 61720031-4bc4-458a-a54c-c18f927f7c26 · outbound

This paper cites Intriguing properties of hyperbolic embeddings in vision- language models,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Intriguing properties of hyperbolic embeddings in vision- language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:28.706939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:27.205584Z digest=sha256:4c4151f05732a449f5e73f8b82ee324c008477abaff9d785d292d768fb3e9bf9

Observation 6ce148ab-c7d8-418f-8e2a-f609ee34b9a6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Learning transferable visual models from natural language supervision,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:54:27.252201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:54:27.252201Z digest=sha256:eeec35ccea86ebc41bdd660159f24fc520cd3743550e1c52bf633d87eb7bc654

Observation e5313521-7af0-41fe-880d-2896a9bbec65 · outbound

This paper cites A multi-view approach to audio-visual speaker verification,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association A multi-view approach to audio-visual speaker verification,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:28.547974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:27.321773Z digest=sha256:ee966d637de5af22639b8b8161377b91a72dbaebcad8106bc4cfa7a382fbbfbc

Observation cb1096bd-df33-4d6a-a6f3-bb053b37cf37 · outbound

This paper cites Adversarial-metric learning for audio-visual cross-modal match- ing,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Adversarial-metric learning for audio-visual cross-modal match- ing,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:28.443312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:27.402068Z digest=sha256:5bd7f27941ab4527559c0f35a12e122800791868bd42b5a3b39fe2323ba5b1dc

Observation 2f57c013-e2fa-4993-b782-ec536781d88a · outbound

This paper cites Disentangled represen- tation learning for cross-modal biometric matching,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Disentangled represen- tation learning for cross-modal biometric matching,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:28.300220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:27.509218Z digest=sha256:4114717341a95f228bf808ce64866a8f10ad7b0d988dcaf16ea7a38edd7f5b02

Observation b1625b69-dfc8-4933-943d-b34dbef6a299 · outbound

This paper cites Gated Multimodal Units for Information Fusion.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Gated Multimodal Units for Information Fusion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:54:27.571280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:54:27.571280Z digest=sha256:8788d2e53f5cb10b963b9a05f00998217121b30abe776c025888f72aaa0292aa

Observation 9ea7d362-61ee-41b7-853c-99a71acc5cec · outbound

This paper cites Deep face recogni- tion,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Deep face recogni- tion,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:28.163281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:27.672172Z digest=sha256:0eb49682cc0b63b443dd93a3506d6a0778079be7274bc14f5e4af1ae69364134

Observation bb7f9a44-e7d4-4fd1-bd9c-efa36bfdf944 · outbound

This paper cites Utterance- level aggregation for speaker recognition in the wild,.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association Utterance- level aggregation for speaker recognition in the wild,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:54:28.007269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:54:27.753449Z digest=sha256:f3e213f325e550c3ca854fcb73e6eb5a785f76392f7769393fcad944c6574d03

Pith citing papers

Observation 19de554e-2010-46a2-95f3-8195cab4a911 · inbound

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association cites this paper.

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:54:25.077226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:54:25.077226Z digest=sha256:d198a34ceead503527872fe07ad4371ccc823ebb51891db1ad69eccaf00f7984

Observation cd2426c6-f426-4597-933d-b22d65292576 · inbound

MuteSwap: Visual-informed Silent Video Identity Conversion cites this paper.

MuteSwap: Visual-informed Silent Video Identity Conversion PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:19:53.919077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.555712Z digest=sha256:885b5d1015e471de946678264e575b8983e2a2e0aa69b42afbf9cd4614790da4