Pith. sign in

Paper Citation Record · LEDGER

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2506.14427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14427 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:21:58.204834Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T15:07:52.715880Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.188094Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact3
  • verified fuzzy43
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3f82b19-44d6-435f-9c02-be0f261f38bb · outbound

This paper cites A review of speaker diarization: Recent advances with deep learning,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A review of speaker diarization: Recent advances with deep learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.659626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:50.470683Z digest=sha256:f355f02db8c169be463a7693ff4069ce2847e7437726e37e631bf1f94442cc72

Observation e13f9fe6-6fd1-4b5a-961a-4aeb8dd64713 · outbound

This paper cites Speaker diarization: A review of recent research,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization: A review of recent research,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.458552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:50.640545Z digest=sha256:a71989d383c2a0871597deae4252c63e4fc7560134ced6be37aaa9422c206717

Observation ec32f24c-0d18-4624-b2e7-4f1c8a84a765 · outbound

This paper cites Speaker diarization with lstm,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization with lstm,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.238820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:50.787115Z digest=sha256:fca408a194a44f460531c8fe7112ff315563e43c52d273d7d0f3d23703d53842

Observation 73ef4eec-b169-4a19-99fd-25a05bff166e · outbound

This paper cites Front- end factor analysis for speaker verification,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Front- end factor analysis for speaker verification,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:50.865767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:50.865767Z digest=sha256:6d56d23ba2dfb7b371fc7274a6581e312fcce2b2146c1ed73d35b3e96cbd6dd4

Observation 9c5e0ead-22dd-4f87-8f96-979565dfbea6 · outbound

This paper cites X- vectors: Robust dnn embeddings for speaker recognition,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset X- vectors: Robust dnn embeddings for speaker recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.974964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:51.002901Z digest=sha256:8e27cdd702e5b8ca4116a7d5c6c1b4388f78401817dcdc76677817b2b74c827a

Observation 56007a15-7a95-4dad-9254-627bcce8099a · outbound

This paper cites Developing on-line speaker diarization system.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Developing on-line speaker diarization system

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.703109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:51.144910Z digest=sha256:faa775a04dc36d5b67a3126e9e5b881f256cf2746c39b2691ade8aba5730373f

Observation 1efab697-e996-43c9-b157-0884261e392a · outbound

This paper cites A study of the cosine distance-based mean shift for telephone speech diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A study of the cosine distance-based mean shift for telephone speech diarization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.304830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:51.314753Z digest=sha256:7cf1e03c65838511a10804e49fcfeda7bb0400fff1a68aff9a6c0aa582ebf5dc

Observation 09e744c1-aa83-4cfa-bc3a-41559e4f6828 · outbound

This paper cites Speaker diarization with lstm,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization with lstm,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:11.883861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:51.517631Z digest=sha256:5a5579721b53272ca72f27507085ac8e7b3dd25dfba38601b1f70736355f7150

Observation fde14acd-eef8-4b76-a3dd-e02491bab023 · outbound

This paper cites A robust stopping criterion for agglomerative hierarchical clustering in a speaker diarization system.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A robust stopping criterion for agglomerative hierarchical clustering in a speaker diarization system

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:11.374824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:51.731444Z digest=sha256:4b1f9bf56ae3d4a363f3cda5dc4cf7521e50b48589fc476440c03fd4e0dc8c4c

Observation ad39687d-62ea-4546-9eb8-9db7c3a523f3 · outbound

This paper cites Characterizing performance of speaker diarization systems on far-field speech using standard methods,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Characterizing performance of speaker diarization systems on far-field speech using standard methods,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.884984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:51.890542Z digest=sha256:4926301f842b91e37ba270e80f0924b24ba70fab2cfaad6046eb92e787342c90

Observation f87ca2c2-33fb-4503-8ebe-fe82f85839e9 · outbound

This paper cites Discriminative neural clustering for speaker diarisation,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Discriminative neural clustering for speaker diarisation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.494758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:52.094815Z digest=sha256:0bc4a5710129dd1bb238f3d4dc32e642ec569af4f5ccc69567bc2601acb7065f

Observation 1799a67a-a6c6-4938-a63d-5af4012ce448 · outbound

This paper cites Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implemen- tation and analysis on standard tasks,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implemen- tation and analysis on standard tasks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.129979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:52.305116Z digest=sha256:78961b859b0f7a07604500ea2f27a6d3d4366ea8d3326530e693977dcbd5ddb5

Observation 6c09ff5f-6367-4f97-b3af-add1a7bd70f9 · outbound

This paper cites End-to-End Neural Speaker Diarization with Permutation-Free Objec- tives,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-End Neural Speaker Diarization with Permutation-Free Objec- tives,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.865817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:52.444837Z digest=sha256:53c90c3002b7842aa6225fea5488da66c3883360b6dfe2e0548dfe93c405a243

Observation 6962df6d-2f1d-46a9-b7ff-7e1488b08cd9 · outbound

This paper cites End-to-end neural speaker diarization with self-attention,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-end neural speaker diarization with self-attention,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:52.593916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:52.593916Z digest=sha256:c5e8a03ae92d294ec31c7f34cc7b53b54f8c6f09e2a0608f9b5d235511282682

Observation 7a66dfec-ffab-4768-bbe7-c47f87b803ec · outbound

This paper cites Auxiliary loss of transformer with residual connection for end-to-end speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Auxiliary loss of transformer with residual connection for end-to-end speaker diarization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.614817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:52.694892Z digest=sha256:fcb3a415f5422bfb66809ed0775fcf071414435a9f889acb74dea4e54d20f593

Observation 91d2595b-6fe6-4cd8-ac5e-ddf332add929 · outbound

This paper cites Incorporating end-to-end framework into target- speaker voice activity detection,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Incorporating end-to-end framework into target- speaker voice activity detection,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.328883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:52.834754Z digest=sha256:a0a15998b300dc8c36d7d35aae497538b7cbb9459699311de0ca457ad0729bb7

Observation 48b1ea25-89ef-42ed-8019-a486ea98782c · outbound

This paper cites Target-Speaker V oice Activity Detection: A Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Target-Speaker V oice Activity Detection: A Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:08.528593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:52.935871Z digest=sha256:58d71f30777bdea93e4b2b0934af9901dc89f556412d60eefb2396cd481e6162

Observation e72957de-844b-4af2-b005-b6d83bdb7240 · outbound

This paper cites Ansd-ma-mse: Adaptive neural speaker diarization using memory-aware multi-speaker embed- ding,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Ansd-ma-mse: Adaptive neural speaker diarization using memory-aware multi-speaker embed- ding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:08.157893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:53.134746Z digest=sha256:95fa16522ddec3b76f212f3283f9c843be458065a12ff217dd1d83b7b5c92bee

Observation 2716b444-1e63-4fc5-9248-731d85914f91 · outbound

This paper cites The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:53.248557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:53.248557Z digest=sha256:1adb0e9ad9cb276ba1783125d703ebe1f649f10a0b1d32ec4a476cc3bd6cb119

Observation 8ff547e6-29f0-42fd-a60f-6f86231dddc2 · outbound

This paper cites The ustc-nercslip systems for chime-7 challenge,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The ustc-nercslip systems for chime-7 challenge,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.773098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:53.427755Z digest=sha256:a7e9fd03c131485757b8c8de582fdaba8b7388705ba1f9b7aebc2837cb96ea08

Observation bca27e2b-c9db-4d31-a9a3-d0314d2ed985 · outbound

This paper cites Audio-visual speaker diarization based on spatiotemporal bayesian fusion,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Audio-visual speaker diarization based on spatiotemporal bayesian fusion,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.537163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:53.571500Z digest=sha256:e38af7c5bc3b6fe9680628623413a6964364408c259f88a87c20cba48447df35

Observation c773421d-5f15-4cde-850b-8864ec10fbf1 · outbound

This paper cites Quantitative association of vocal-tract and facial behavior,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Quantitative association of vocal-tract and facial behavior,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.309079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:53.778190Z digest=sha256:07278228830d2847e7de85baa6f783a35f21c42b714b8e84c36a7701721119aa

Observation e43a8855-a018-4c74-aa81-c275a07f6cf7 · outbound

This paper cites Multimodal speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Multimodal speaker diarization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.064881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:53.934772Z digest=sha256:fe36485751fc7bd10545bb17f9889ef92f3738e27c7e0f60f0e17c8d24f123f5

Observation ce9ae35e-a165-40e7-985d-40c08a73cb89 · outbound

This paper cites Who said that?: Audio-visual speaker diarisation of real-world meetings,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Who said that?: Audio-visual speaker diarisation of real-world meetings,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:06.764776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:54.094834Z digest=sha256:f44cb723282b70d4a60048268233f9e5b1dcba4e1894aedaa7876cf4b445a7a0

Observation 2e95f8fa-c44d-44f1-87d8-08440b71e7b2 · outbound

This paper cites Spot the conversation: Speaker diarisation in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Spot the conversation: Speaker diarisation in the wild,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:06.436582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:54.254227Z digest=sha256:f687850d15aa5b05931a5f2354b64a2b2baa1535b1e258a2f49c0eba96463485

Observation 428b31c9-85f8-4f23-9c29-9dac8e0da984 · outbound

This paper cites End-to-end audio-visual neural speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-end audio-visual neural speaker diarization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.908267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:54.571101Z digest=sha256:0f674befe6ada44ac0e7207bd905c4ddbc137012e6184634586fb0b3c9e67fb0

Observation c140ea81-7ae9-40be-b600-f63c29793861 · outbound

This paper cites Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:59.726773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:54.869454Z digest=sha256:6ee3c70450d95f131e0ce7f8721e85fdce28cb79fc120e0188c531b7949d15ab

Observation 0fa53158-fd14-451c-907c-60b3a82249c7 · outbound

This paper cites Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:59.396178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:55.024954Z digest=sha256:8ed50ba50b9b8fde27862cc88023790c94e4ce87647179004e08165039480920

Observation 2ffe891a-1a73-45a7-907e-004876e43b21 · outbound

This paper cites Semi-supervised multi-channel speaker diarization with cross- channel attention,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Semi-supervised multi-channel speaker diarization with cross- channel attention,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.602331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:55.124886Z digest=sha256:56482d6e1f423f6e2400cb69569d4a1e11ad6dd38f3dbf983e138b66555dd4af

Observation e1c36b09-1a9e-4a1b-8abf-6cfe9b2aafed · outbound

This paper cites Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.349237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:55.207393Z digest=sha256:4bbda251d7e40d9caf9b55b586319316385106364736d5fd6ffed75231f68a33

Observation 026db9a8-134e-4b29-b845-835258614c2c · outbound

This paper cites Ava-avd: Audio-visual speaker diarization in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Ava-avd: Audio-visual speaker diarization in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.115533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:55.324692Z digest=sha256:7277f2155579b8d34661e351a72406e917a57e02b42e52cefd50da44b8ae3003

Observation 05fbe6af-a073-45db-bb76-3551a48e1ab3 · outbound

This paper cites Semi-supervised training with pseudo-labeling for end- to-end neural diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Semi-supervised training with pseudo-labeling for end- to-end neural diarization,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:04.756713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:55.483931Z digest=sha256:4706f84b925e8e8e1a8ca1499e64c4a331f12a49bcf4c20fc47b1588a824e645

Observation f4fd928c-abbd-4ec0-8ea1-5b27ae7b4a26 · outbound

This paper cites Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:04.309737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:55.624749Z digest=sha256:36450b80eacbe35515a7d3e0a8ce38da6291f351b178c54de2edd8e2197b3d2a

Observation ac8da1ae-9ff5-4ccc-b653-e77e25873ae7 · outbound

This paper cites The multimodal information based speech processing (misp) 2022 challenge: Audio- visual diarization and recognition,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The multimodal information based speech processing (misp) 2022 challenge: Audio- visual diarization and recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.958320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:55.784828Z digest=sha256:bbdec9f11420100a777811c14ce2c5a4e91a8326d889c7c0b57c51e3042f1531

Observation 284f3f73-a188-42a9-86b4-331cfea318fa · outbound

This paper cites Notsofar-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Notsofar-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.664897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:55.902011Z digest=sha256:aa0f22e1cac2b3f39e6ca1025452d11e081ce90486756d5a7aa725f23b6e547e

Observation cedac64d-e72b-4cea-976f-faf267d75d7e · outbound

This paper cites The nist speaker recognition evaluations: 1996-2001.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The nist speaker recognition evaluations: 1996-2001

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.410058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:56.047128Z digest=sha256:736f22a231e089ed8fd134c6431e7b066ff05dc5134cbb3eaed7e572cd674cb4

Observation 6dad5101-8ed9-4e39-b33c-813a367f287b · outbound

This paper cites First dihard challenge evaluation plan,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset First dihard challenge evaluation plan,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.186514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:56.162241Z digest=sha256:cb9ddeca12e7ac6b167602fd7da491f1df1a3ad41bcac3bd5a3d187f31c40eca

Observation 68cf3cf7-1053-4eba-8bcb-860721340f5e · outbound

This paper cites The Second DIHARD Diarization Challenge: Dataset, task, and baselines.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Second DIHARD Diarization Challenge: Dataset, task, and baselines

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:58.984753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:56.288230Z digest=sha256:277ace6d79caab161efb7303e901e171911c510ea4b82aaa04e8384f74c18596

Observation 7b1e1bb5-efde-4eee-9410-07ad6e9434eb · outbound

This paper cites The Third DIHARD Diarization Challenge.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Third DIHARD Diarization Challenge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.433967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.433967Z digest=sha256:532e4ce48e6901f5fa6580bc125dd0a51a9e101ad9865536935c4d27590ea5a1

Observation 5daffbd3-c75b-4dee-adf6-aea506583848 · outbound

This paper cites M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.864740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:56.562953Z digest=sha256:8d9e9b6bac49f621db7cb0e2d01199902717fad3bb21abe4ad4186d4360ec5a3

Observation f88caf64-7715-4f24-94b6-681e862c35be · outbound

This paper cites The ami meeting corpus,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The ami meeting corpus,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.615362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:56.667720Z digest=sha256:6f18dcfe2bb3bf791be410741a6fd28ea54b6250d0f5bcc6facdcae36682acf3

Observation dca6c443-3407-4b0b-92fb-977eb884cbde · outbound

This paper cites The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.777201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.777201Z digest=sha256:623398c14538e0ead60c50f767bcbfb2c05db6a2ab0e7b721c041fed0db25911

Observation 5438a602-d4a0-4270-b8f1-af5dabe50ec5 · outbound

This paper cites Msdwild: Multi-modal speaker diarization dataset in the wild.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Msdwild: Multi-modal speaker diarization dataset in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.314881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:56.893682Z digest=sha256:1956412fb2c44897f57cbfd89d2e9f51912e9138cb7dc3fd923c43a44a7ff555

Observation 17e3e495-23e7-4b84-b33e-6598d3733f73 · outbound

This paper cites An approach to scene change detection,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset An approach to scene change detection,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.056694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:56.984747Z digest=sha256:6338901ffbe39b987528b8658b611b715a9ffb7384f061f1c81f67dc8c90bf3f

Observation d947b473-73bc-4c55-b50d-545945fa1c29 · outbound

This paper cites Dnsmos p. 835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Dnsmos p. 835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.804829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:57.111910Z digest=sha256:a52f6faab937f7ecd232183a82745b488aae676fc3430717cadd2282cbe38525

Observation b6875ee2-91e5-4828-9401-ffd2f0c4e60b · outbound

This paper cites Md-vqa: Multi-dimensional quality assessment for ugc live videos,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Md-vqa: Multi-dimensional quality assessment for ugc live videos,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.506885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:57.220765Z digest=sha256:2bdcfcb38ec7cbde3a1c4c937bb343c4ec28416249f81f8136c9445529748e26

Observation c6a24fa6-3ee0-42bd-a628-9d00325f3f8b · outbound

This paper cites Out of time: automated lip sync in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Out of time: automated lip sync in the wild,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.255798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:57.380226Z digest=sha256:d16863081af4208b422a3c1ef437b804919fff6fc073c4405a5f3f98b9933794

Observation 8a3154c0-925e-43c5-bf15-519b2d0f01ca · outbound

This paper cites Retinaface: Single-shot multi-level face localisation in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Retinaface: Single-shot multi-level face localisation in the wild,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.979743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:57.559593Z digest=sha256:7283394b0624859f14181505d8d14a343b5259c7168937b6aacd6a68389326f4

Observation 6c5c5a94-fc8e-4f58-be3c-5990295ce602 · outbound

This paper cites Simple online and realtime tracking with a deep association metric,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Simple online and realtime tracking with a deep association metric,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.738802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:57.702925Z digest=sha256:7371b0ddd394b3f3e02ad6cee190a37598b33588b1a64d1abe010416965685d4

Observation a6efce1f-5493-45e2-acdb-574c78d25f91 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset MediaPipe: A Framework for Building Perception Pipelines

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:57.834750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:57.834750Z digest=sha256:47322d5f852a96e4703d7e03c7776d8bff90cd1900ebcd502c67114c4ca6baf8

Observation 7e4dfd2a-1d75-4bae-a506-09f2c1823712 · outbound

This paper cites 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.464756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:58.044836Z digest=sha256:1c46401824cacb49febd508e1cbb2bbb2e35133228ade90336bcd520d99dad73

Observation 10058e43-5b10-47f4-b6a0-c3e06c096d82 · outbound

This paper cites Dover-lap: A method for combining overlap- aware diarization outputs,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Dover-lap: A method for combining overlap- aware diarization outputs,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.209099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:21:58.204834Z digest=sha256:404f35219d6c678a9e995c53f1c7b46149f5e2ca00bf453f30bec0b5762ded07

Pith citing papers

Observation 627f9364-7480-480c-a8e3-e381cc071047 · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.189364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:98844e220e1b018d7ea82fd08a491570b7805971e85e61f29f93ac43d6941db5