Pith. sign in

Paper Citation Record · LEDGER

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

As of 22 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2506.14427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14427 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:21:58.204834Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T15:07:52.715880Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.188094Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact3
  • verified fuzzy43
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3f82b19-44d6-435f-9c02-be0f261f38bb · outbound

This paper cites A review of speaker diarization: Recent advances with deep learning,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A review of speaker diarization: Recent advances with deep learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.659626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:50.470683Z digest=sha256:88329a8e25c297f6f45de638ee0c53005fee870ada7a946ac616534c7595654a

Observation e13f9fe6-6fd1-4b5a-961a-4aeb8dd64713 · outbound

This paper cites Speaker diarization: A review of recent research,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization: A review of recent research,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.458552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:50.640545Z digest=sha256:9ff7b7517eaaac86e8417e1f828d515ba08e83a89fefe92c6a1d94505b54f21f

Observation ec32f24c-0d18-4624-b2e7-4f1c8a84a765 · outbound

This paper cites Speaker diarization with lstm,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization with lstm,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.238820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:50.787115Z digest=sha256:f9bcb8df1734de20cf45807648c553fbfab4469120979649e9e08fda20b63baa

Observation 73ef4eec-b169-4a19-99fd-25a05bff166e · outbound

This paper cites Front- end factor analysis for speaker verification,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Front- end factor analysis for speaker verification,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:50.865767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:50.865767Z digest=sha256:b3a0d13e9ac77ce8e1e1f4633c9f783a4071a24771543d7b434ef490f97c6700

Observation 9c5e0ead-22dd-4f87-8f96-979565dfbea6 · outbound

This paper cites X- vectors: Robust dnn embeddings for speaker recognition,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset X- vectors: Robust dnn embeddings for speaker recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.974964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:51.002901Z digest=sha256:e2ab07a56d25e4ec4027b4eecc64303ba434fe1dcf49e6e7c00f76d877a5409e

Observation 56007a15-7a95-4dad-9254-627bcce8099a · outbound

This paper cites Developing on-line speaker diarization system.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Developing on-line speaker diarization system

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.703109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:51.144910Z digest=sha256:6589955530f911bad7af39f384bba89b1d37d42ad072ffbf093f5ed7fa6a2b88

Observation 1efab697-e996-43c9-b157-0884261e392a · outbound

This paper cites A study of the cosine distance-based mean shift for telephone speech diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A study of the cosine distance-based mean shift for telephone speech diarization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.304830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:51.314753Z digest=sha256:5b6a3bbbf5fd0c7aec01ba1c669067c60e625519357dbe2aece133869f27ecc5

Observation 09e744c1-aa83-4cfa-bc3a-41559e4f6828 · outbound

This paper cites Speaker diarization with lstm,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization with lstm,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:11.883861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:51.517631Z digest=sha256:1b86ff7707c6b436924ded74d740e98aceaba0000647f8084b42b35a15d329c3

Observation fde14acd-eef8-4b76-a3dd-e02491bab023 · outbound

This paper cites A robust stopping criterion for agglomerative hierarchical clustering in a speaker diarization system.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A robust stopping criterion for agglomerative hierarchical clustering in a speaker diarization system

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:11.374824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:51.731444Z digest=sha256:a3b7b656677be92853c9ccd58ef920f84cb75d11a999e52956da30ba3390be16

Observation ad39687d-62ea-4546-9eb8-9db7c3a523f3 · outbound

This paper cites Characterizing performance of speaker diarization systems on far-field speech using standard methods,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Characterizing performance of speaker diarization systems on far-field speech using standard methods,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.884984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:51.890542Z digest=sha256:bd546065d473080b0ceccddbb3582f94829d198ecc78374f77e07d7d82a6a844

Observation f87ca2c2-33fb-4503-8ebe-fe82f85839e9 · outbound

This paper cites Discriminative neural clustering for speaker diarisation,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Discriminative neural clustering for speaker diarisation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.494758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:52.094815Z digest=sha256:761c969dde01891ac12a91459d2f70b735aa10c3aa7004881a2bdd59e7654959

Observation 1799a67a-a6c6-4938-a63d-5af4012ce448 · outbound

This paper cites Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implemen- tation and analysis on standard tasks,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implemen- tation and analysis on standard tasks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.129979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:52.305116Z digest=sha256:6ea442ffe4034c30e60591b9812546d069018844d9f95501be0aaa5c91a77079

Observation 6c09ff5f-6367-4f97-b3af-add1a7bd70f9 · outbound

This paper cites End-to-End Neural Speaker Diarization with Permutation-Free Objec- tives,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-End Neural Speaker Diarization with Permutation-Free Objec- tives,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.865817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:52.444837Z digest=sha256:47e1b17e2cc001e5e76a4716dfcf502c768fce10e745169e65a1f314a46612fc

Observation 6962df6d-2f1d-46a9-b7ff-7e1488b08cd9 · outbound

This paper cites End-to-end neural speaker diarization with self-attention,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-end neural speaker diarization with self-attention,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:52.593916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:52.593916Z digest=sha256:de710c2a3d3bbbec8214f849b0d659acd0d637e1113feedb127a25c2f5fc2520

Observation 7a66dfec-ffab-4768-bbe7-c47f87b803ec · outbound

This paper cites Auxiliary loss of transformer with residual connection for end-to-end speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Auxiliary loss of transformer with residual connection for end-to-end speaker diarization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.614817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:52.694892Z digest=sha256:f2207e9ae00da4d52d9bf2670184dd5609109cea50b8bd1ed11139890675f070

Observation 91d2595b-6fe6-4cd8-ac5e-ddf332add929 · outbound

This paper cites Incorporating end-to-end framework into target- speaker voice activity detection,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Incorporating end-to-end framework into target- speaker voice activity detection,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.328883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:52.834754Z digest=sha256:27450fe21c208308247bbcb51faeca6eb958a98f6e266e166c275e900f779111

Observation 48b1ea25-89ef-42ed-8019-a486ea98782c · outbound

This paper cites Target-Speaker V oice Activity Detection: A Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Target-Speaker V oice Activity Detection: A Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:08.528593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:52.935871Z digest=sha256:14ab58de8111436cba3c0a6b30b7687a1deb4a89fdef4511bec47dbdd6cff800

Observation e72957de-844b-4af2-b005-b6d83bdb7240 · outbound

This paper cites Ansd-ma-mse: Adaptive neural speaker diarization using memory-aware multi-speaker embed- ding,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Ansd-ma-mse: Adaptive neural speaker diarization using memory-aware multi-speaker embed- ding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:08.157893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:53.134746Z digest=sha256:03116f80e8d842cbcab4517af3542f34f51c08429c4d98bced6326f21c468867

Observation 2716b444-1e63-4fc5-9248-731d85914f91 · outbound

This paper cites The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:53.248557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:53.248557Z digest=sha256:96976fc35e6f13365b2df955d09c54b79bdf489cecca2cf81d51fb5200786913

Observation 8ff547e6-29f0-42fd-a60f-6f86231dddc2 · outbound

This paper cites The ustc-nercslip systems for chime-7 challenge,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The ustc-nercslip systems for chime-7 challenge,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.773098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:53.427755Z digest=sha256:ea39b23432b5b39c9edff65fdfb55af5819dacfad7222991ae4313912ba12687

Observation bca27e2b-c9db-4d31-a9a3-d0314d2ed985 · outbound

This paper cites Audio-visual speaker diarization based on spatiotemporal bayesian fusion,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Audio-visual speaker diarization based on spatiotemporal bayesian fusion,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.537163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:53.571500Z digest=sha256:2a03e58b246614964797c8dcbcaab7f62e98dc1b1b418c084be2cac1e877396e

Observation c773421d-5f15-4cde-850b-8864ec10fbf1 · outbound

This paper cites Quantitative association of vocal-tract and facial behavior,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Quantitative association of vocal-tract and facial behavior,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.309079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:53.778190Z digest=sha256:aeabf9a6a006bb3e2f47c8bbe1181e9a822250d5d5526610b1ed56c1a1be8e1a

Observation e43a8855-a018-4c74-aa81-c275a07f6cf7 · outbound

This paper cites Multimodal speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Multimodal speaker diarization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.064881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:53.934772Z digest=sha256:d685a1fe260ede164cc83a88fee29aed06e9b4379f02e00c4995700c26e84b70

Observation ce9ae35e-a165-40e7-985d-40c08a73cb89 · outbound

This paper cites Who said that?: Audio-visual speaker diarisation of real-world meetings,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Who said that?: Audio-visual speaker diarisation of real-world meetings,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:06.764776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:54.094834Z digest=sha256:b12719b487ce4b6797bf4bdbed6e0834a1c9f25d9ea8c518d2f56407f87fc21c

Observation 2e95f8fa-c44d-44f1-87d8-08440b71e7b2 · outbound

This paper cites Spot the conversation: Speaker diarisation in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Spot the conversation: Speaker diarisation in the wild,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:06.436582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:54.254227Z digest=sha256:85b6cb87238890618c77b6ef01b30f2a2389bd8cefd6838683d1b1ec9fa3c954

Observation 428b31c9-85f8-4f23-9c29-9dac8e0da984 · outbound

This paper cites End-to-end audio-visual neural speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-end audio-visual neural speaker diarization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.908267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:54.571101Z digest=sha256:b67d82dc9fa0b370bbbb2063a69f125ed16ddb9f43b5ced3aa62a67777cb469c

Observation c140ea81-7ae9-40be-b600-f63c29793861 · outbound

This paper cites Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:59.726773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:54.869454Z digest=sha256:f137b7ec0c2955435726fee5868d6e83ba2405c4815ebdf55114c3f5341d4fc6

Observation 0fa53158-fd14-451c-907c-60b3a82249c7 · outbound

This paper cites Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:59.396178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:55.024954Z digest=sha256:8b1b005fc04f590b619d36d465c5b39678aeb435e502deab124895e64ad21a71

Observation 2ffe891a-1a73-45a7-907e-004876e43b21 · outbound

This paper cites Semi-supervised multi-channel speaker diarization with cross- channel attention,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Semi-supervised multi-channel speaker diarization with cross- channel attention,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.602331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:55.124886Z digest=sha256:e79a11012f0f784f86c4de17c0294d744787ab4196d19098c9f823379d45df4a

Observation e1c36b09-1a9e-4a1b-8abf-6cfe9b2aafed · outbound

This paper cites Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.349237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:55.207393Z digest=sha256:56883e9b2bd764484b620322cfe42e47b81a79752e9debc121f6ba7a7204aeb6

Observation 026db9a8-134e-4b29-b845-835258614c2c · outbound

This paper cites Ava-avd: Audio-visual speaker diarization in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Ava-avd: Audio-visual speaker diarization in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.115533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:55.324692Z digest=sha256:a838d21bfd448fb5b842cab8e651bc4785ac73caff96e64e6262469d97dadbbb

Observation 05fbe6af-a073-45db-bb76-3551a48e1ab3 · outbound

This paper cites Semi-supervised training with pseudo-labeling for end- to-end neural diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Semi-supervised training with pseudo-labeling for end- to-end neural diarization,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:04.756713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:55.483931Z digest=sha256:78c82e07aa8b96d7d8169f433f5a7c21d95de11f621dcc81bc37f38c5ff07a01

Observation f4fd928c-abbd-4ec0-8ea1-5b27ae7b4a26 · outbound

This paper cites Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:04.309737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:55.624749Z digest=sha256:1abdfa8ed7377fbc834e3063e407afbdcf254655d542905a37df50e232391d25

Observation ac8da1ae-9ff5-4ccc-b653-e77e25873ae7 · outbound

This paper cites The multimodal information based speech processing (misp) 2022 challenge: Audio- visual diarization and recognition,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The multimodal information based speech processing (misp) 2022 challenge: Audio- visual diarization and recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.958320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:55.784828Z digest=sha256:22f5f624a3011f6316288b8b20ab65bd0df05aebd9d822b376777230b52e7385

Observation 284f3f73-a188-42a9-86b4-331cfea318fa · outbound

This paper cites Notsofar-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Notsofar-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.664897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:55.902011Z digest=sha256:7365cbeb9f2d1a37e77944245b4b4974eafcd332bb009fbd2687c331a980eabf

Observation cedac64d-e72b-4cea-976f-faf267d75d7e · outbound

This paper cites The nist speaker recognition evaluations: 1996-2001.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The nist speaker recognition evaluations: 1996-2001

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.410058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:56.047128Z digest=sha256:bbbf953b48e190a589901e17d54d1038ab20fcce8834cdd21e1ff9d37174da5a

Observation 6dad5101-8ed9-4e39-b33c-813a367f287b · outbound

This paper cites First dihard challenge evaluation plan,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset First dihard challenge evaluation plan,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.186514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:56.162241Z digest=sha256:37c2fcd52f83efdbe8d90157a18e52a3be78ac0973eb175050f4947d29651475

Observation 68cf3cf7-1053-4eba-8bcb-860721340f5e · outbound

This paper cites The Second DIHARD Diarization Challenge: Dataset, task, and baselines.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Second DIHARD Diarization Challenge: Dataset, task, and baselines

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:58.984753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:56.288230Z digest=sha256:da0ae65414754e98cdc22fdb8c6edc72175c78b0dcab3e58371b4481d24769a8

Observation 7b1e1bb5-efde-4eee-9410-07ad6e9434eb · outbound

This paper cites The Third DIHARD Diarization Challenge.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Third DIHARD Diarization Challenge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.433967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.433967Z digest=sha256:62bce78c38e05da0ff188edfa7e8ebbbf197b9d3a6c4ab6df5636a577dce1a54

Observation 5daffbd3-c75b-4dee-adf6-aea506583848 · outbound

This paper cites M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.864740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:56.562953Z digest=sha256:682ad387676ccdb656fdf4351daeaaae4465d36c076db16329f2bd4819605071

Observation f88caf64-7715-4f24-94b6-681e862c35be · outbound

This paper cites The ami meeting corpus,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The ami meeting corpus,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.615362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:56.667720Z digest=sha256:a7479c8389a3d5b47c55840b9e74c7b12184b648d51d2e003b94729d4a5768e7

Observation dca6c443-3407-4b0b-92fb-977eb884cbde · outbound

This paper cites The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.777201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.777201Z digest=sha256:329d269e7d5b38f3805c7852c11fda2f3e2f42724943121ade75142fa723130c

Observation 5438a602-d4a0-4270-b8f1-af5dabe50ec5 · outbound

This paper cites Msdwild: Multi-modal speaker diarization dataset in the wild.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Msdwild: Multi-modal speaker diarization dataset in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.314881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:56.893682Z digest=sha256:35dd727aa33df8b365c61ade02ed93c9fe6367d775198ec56855a2dd1f2ab832

Observation 17e3e495-23e7-4b84-b33e-6598d3733f73 · outbound

This paper cites An approach to scene change detection,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset An approach to scene change detection,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.056694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:56.984747Z digest=sha256:f9426418bad6f3ff17a6aa97ce098849fecb23b54e593acc9673d0d3b4e76123

Observation d947b473-73bc-4c55-b50d-545945fa1c29 · outbound

This paper cites Dnsmos p. 835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Dnsmos p. 835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.804829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:57.111910Z digest=sha256:3ee76c796ce32f4893faa4c63dedf4b78cc77c1888305a0b597bd014b63c9065

Observation b6875ee2-91e5-4828-9401-ffd2f0c4e60b · outbound

This paper cites Md-vqa: Multi-dimensional quality assessment for ugc live videos,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Md-vqa: Multi-dimensional quality assessment for ugc live videos,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.506885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:57.220765Z digest=sha256:b455438e26b13e3fed976a147838ecb44fa2dccfa40f78f2baa22956658ea04f

Observation c6a24fa6-3ee0-42bd-a628-9d00325f3f8b · outbound

This paper cites Out of time: automated lip sync in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Out of time: automated lip sync in the wild,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.255798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:57.380226Z digest=sha256:f0f67aad0eafb705c3afe663b355edbd99d116e175fa1a7107ddbf63fa4e7ac4

Observation 8a3154c0-925e-43c5-bf15-519b2d0f01ca · outbound

This paper cites Retinaface: Single-shot multi-level face localisation in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Retinaface: Single-shot multi-level face localisation in the wild,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.979743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:57.559593Z digest=sha256:a606e338b2d6225233211cdbd104568cfad14f299bcfba6fdbdc7a68a5ccb293

Observation 6c5c5a94-fc8e-4f58-be3c-5990295ce602 · outbound

This paper cites Simple online and realtime tracking with a deep association metric,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Simple online and realtime tracking with a deep association metric,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.738802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:57.702925Z digest=sha256:9d0ce5d235c965a0e18e57bf95d80c4f4c8fec5973a78701a42b967f869605b6

Observation a6efce1f-5493-45e2-acdb-574c78d25f91 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset MediaPipe: A Framework for Building Perception Pipelines

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:57.834750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:57.834750Z digest=sha256:ac4797ffef9fabe0ff790dad48e274c0a9cdc90cf4e4b2ac326fc2ce46803b17

Observation 7e4dfd2a-1d75-4bae-a506-09f2c1823712 · outbound

This paper cites 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.464756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:58.044836Z digest=sha256:156ab55c966ee79be2529485bd3741da6a39fafd51531019661e33a0012af130

Observation 10058e43-5b10-47f4-b6a0-c3e06c096d82 · outbound

This paper cites Dover-lap: A method for combining overlap- aware diarization outputs,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Dover-lap: A method for combining overlap- aware diarization outputs,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.209099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:21:58.204834Z digest=sha256:79ae1bf0a52751cae63f4921742f7b8a2c34f5af16e6c97d13c526f4a83846fd

Pith citing papers

Observation 627f9364-7480-480c-a8e3-e381cc071047 · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.189364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:1f66cf2465112aafea7b5cc2641ca8405de757e500c81be3535158cce490919f