Pith. sign in

Paper Citation Record · LEDGER

Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2201.02184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.02184 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:44.345646Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.728084Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e7a1966a-fbb3-4274-88fc-a3c1116e54fb · inbound

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition cites this paper.

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:44.345646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:44.345646Z digest=sha256:b2d11a155bfbff6b1a29bd491a94bcb51fa9e97e4bca2e75516f13fec0e43463

Observation fd50aea4-84af-4bc9-8022-d12eb1547d04 · inbound

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models cites this paper.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:50.391701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:50.391701Z digest=sha256:e015d141b67043b380a16fded786e88e9149495e6a295d97c9585803145a3627

Observation 3fcc39cd-9f1e-49b8-8764-0adcdd1345be · inbound

MuteSwap: Visual-informed Silent Video Identity Conversion cites this paper.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.685389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.685389Z digest=sha256:3a6716c8d552cd391db9b47615b03eece6d7a502a9ab509dc01dd49bec5cf51e

Observation 120056b9-0ef1-4f67-994f-127feb7cfcbe · inbound

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring cites this paper.

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:14:35.054556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:14:35.054556Z digest=sha256:5bb238998e4c698f4cdd8d235b7ede6fbc74068b8cf9a5529c587bcf72e488d5

Observation 808bfc5e-bf43-4fc8-9399-8e9f2e1dc2de · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.745331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.745331Z digest=sha256:179dda6e81e8076130edc09d32adb14b7132cdbc58c2ca9a801f01323d83b4a7

Observation e9117a31-9416-46fa-9714-2941e7c61f58 · inbound

HumanOmni-Speaker: Identifying Who said What and When cites this paper.

HumanOmni-Speaker: Identifying Who said What and When Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.524502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:57:12.239390Z digest=sha256:73696f1817a6a9ce25666adbd88ef6fbda37237348cc80a21e4f78096b9414b4

Observation 16998ea3-2045-49cb-80c7-b9342ea01be9 · inbound

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge cites this paper.

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:29.101624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T07:49:44.117839Z digest=sha256:bb15d1edf2824942ce5c42a34f082aaa82719543d58e90323511dc3f8fbec99d

Observation a8176aad-ec4b-4cb7-b5af-e25c1a36bbb5 · inbound

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge cites this paper.

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T19:21:19.583012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:21:19.583012Z digest=sha256:be42f4a37019e342241d3369e538924430a899956cfe31b72a95006d6ab679b0

Observation 709bdce4-0dfc-4a2d-b7a0-eb7505c97883 · inbound

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework cites this paper.

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:30.417122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:23:27.355803Z digest=sha256:7f1ff414f647c175a0dac35eca194551bd800bb37e236514ee2e413e2958f9ad

Observation caa96176-d73b-41ad-9fca-42357ca0e1d2 · inbound

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization cites this paper.

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:20:24.482019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:19:38.661190Z digest=sha256:f8851e8e5eac0232bcef6fcdcecafe678929d2602b08f424c8c53e23b4d1caa2

Observation c1ac8eb4-36f1-4f64-a1c0-65dcd4a3abf9 · inbound

Your Multimodal Speech Model Says I Have a Face for Radio cites this paper.

Your Multimodal Speech Model Says I Have a Face for Radio Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:53:14.062161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:44:50.169335Z digest=sha256:bb74d8d45e76d652242e74e8be5e7e58eeebf36651d9de7432d2767026576fdc

Observation 994b35a5-0354-4e72-b202-f8ccd82040fe · inbound

Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography cites this paper.

Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:47:28.354383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:49:19.643889Z digest=sha256:e5c78df6a1ac80dfdcca18d699099f0b4345f0395502c1e898519c36812d0b05

Observation 641792ed-feaf-437a-8aec-8387dd5eddc8 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.642562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:4d2a4cfefd97101ce0c4e4fdf311fd1f267cf34f8b31e8283b5377452c3e7aee

Observation 68ffcd41-0334-43b0-9014-0b48bc71b45c · inbound

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading cites this paper.

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:47:35.216339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T14:54:08.119264Z digest=sha256:df313193b9d462b10677a74c7698ca81a44ad96982c4ceee9037773cc2af621c

Observation e54e3b54-3554-46c1-b3f0-f9ec27863513 · inbound

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning cites this paper.

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.729646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T11:03:05.449385Z digest=sha256:76242f6c93debb4b5943095cd9f8eee89a3adfcb3ff5e7550f6d4c05932468cb

Observation 524a6f4f-0fc6-4ba4-a4ca-47e4a843d284 · inbound

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement cites this paper.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:49:01.368464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b66e6cf70b935b408cd2617bb3c6979f59d347babe175fc6ff87a86528b4e533

Observation 47d8d709-c8a7-4fa2-9f76-942b31511c2b · inbound

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection cites this paper.

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T02:09:52.116351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:09:52.116351Z digest=sha256:140b5fa5882164e41bb671e34ec86482d4d270aa02877fc4095994becd77b828

Observation 00cdfd15-4ba4-4c9b-bfe7-2a6d6267fb59 · inbound

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE cites this paper.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:03.514200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:02:03.514200Z digest=sha256:7599870db9b398b364fc596d98dae643c7c836214aaaaa3b1e3246c0c7527e5e