Pith. sign in

Paper Citation Record · LEDGER

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2505.13971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13971 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:05.021485Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:00.946933Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T08:09:51.394586Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d516a4d6-1b5c-43fa-bd6f-344aae0d6d84 · outbound

This paper cites The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:00.946933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:00.946933Z digest=sha256:d028ef62bc787f23aac46c1fd7feafa65dfbc5e248040002f9019307a2bc2c83

Observation d182b9a8-a55d-4e53-a58c-b0807a475f2c · outbound

This paper cites Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:10.405861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.015198Z digest=sha256:a5bc569ff71fe7a30bcc82143281b4d1c86c34188235e04a4ceaa5c423f12ab7

Observation 4617894b-5ab6-4b5d-b4c4-abceef6f4b8f · outbound

This paper cites who spoke when.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition who spoke when

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:10.171540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.131980Z digest=sha256:dbb76d12cb4dd86dee325c133d0fd1c6e831ddb7b1f56fcc3d43199b9c2fc4ca

Observation 63d2ca25-9035-4f8f-9644-6e043c78bd31 · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:10.001010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.253337Z digest=sha256:79d6b5bda0ec36aaca3c942e601e9cb99a190b6fe36e56bc638202a50a7f5ac9

Observation af8c5e2b-7ce1-461d-9630-7081a33c5d9e · outbound

This paper cites (2), whereNspk is the total number of speakers in the session.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition (2), whereNspk is the total number of speakers in the session

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.796298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.356018Z digest=sha256:4409d96d2c6eca5acacfdb20aaef9147550bab56d55d57a348fb86d1b31756b1

Observation fd03b07a-e841-49f1-965a-560307c032da · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:09.545159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.496388Z digest=sha256:a464f90580e4eef9c46bf61bae44e17ee595c3976cb91cd59cba3dd1aca29252

Observation fabb067c-e0b7-4de2-9160-4bc51668a496 · outbound

This paper cites A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.281598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.616838Z digest=sha256:4606964cf0af505e659f09e9dc138c6a303eb316916d1632331605db69d7fc82

Observation 095c7367-1392-4e02-a4bd-1df23564b5f2 · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:09.111934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.738485Z digest=sha256:a5894e842c8823146f8f3936bbccd31a99a3dc7f446921242202d1df028889a9

Observation 1ee43906-d77d-4382-a9c1-11056bb6132d · outbound

This paper cites The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.828947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.825645Z digest=sha256:7d051c200b2027164f4e28ccdc740c8a8a675fa088a9002446cb2e10f52d06a7

Observation 94b03a91-ce32-48d7-8791-45fde30cb4a5 · outbound

This paper cites M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.634867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:01.970675Z digest=sha256:368f94a51b2f77143275d4f69ae8e04f0c51c97e178375279d2c004e2e33757a

Observation 474dda51-9528-4bbb-86e4-476a84c651c5 · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.079339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.079339Z digest=sha256:32bd4c7025cacba024e764482769cd80a6f01e50a48b404b7afd422334166771

Observation 24dda821-8acf-4170-af19-ffb35fec1c65 · outbound

This paper cites Continuous speech separation: Dataset and analysis,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Continuous speech separation: Dataset and analysis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.436561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:02.211446Z digest=sha256:47c649339b558279982c8ccbbaecd5ac5dc908fd8759eaa2243aa5a629471757

Observation e9c3e128-15c4-4b6e-ac08-f377e39382e2 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Lib- rispeech: an asr corpus based on public domain audio books,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.305820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.305820Z digest=sha256:d6e4897a7568fe5a293d3bc2a449239209a03cf47dc9cc89ef5354fdd8ceb7fd

Observation ef7adc73-dd1c-4ab8-b3d1-cf8b99dd345f · outbound

This paper cites CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.186566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:02.459989Z digest=sha256:07e951c1965947031c76ce3ebacf121ce965c712b5c1f153dce27e71a8eaea49

Observation 5b477c2d-8e4a-450f-bd4f-280b2da4535e · outbound

This paper cites The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.989009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:02.614902Z digest=sha256:5d0f022380e5ec70d8844e2330c531617b942a11bd320da16e8f86758d4575d9

Observation bea85d7f-dd75-41a5-bfda-ee982ef8a884 · outbound

This paper cites The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.797124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.797124Z digest=sha256:f19ee2731dad3cdf163c925aef04c9969438d87f208b7eae854b0520658e4b09

Observation 9c480d17-e3da-41ee-96c7-d54cb0103dbc · outbound

This paper cites The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.764255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:02.876803Z digest=sha256:2ef1d1ef17abb9178d07335dabedc8e34bb4aa85c0f42df6042a8d04d64bf740

Observation c4e4131e-df16-4fd3-abab-13d0adc854ab · outbound

This paper cites MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.014747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.014747Z digest=sha256:086c4b21f0a16b0cddd759648e1478d277eeee86980aa6b95baa39968f79e088

Observation b75b325c-4bf8-4d38-a95a-702aac5698b2 · outbound

This paper cites End-to-end audio-visual neural speaker diarization,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition End-to-end audio-visual neural speaker diarization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.550360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:03.076635Z digest=sha256:3c40b550faf7060d22ab7e5dadbe2ba4d8892218908358438152b64dc1282f77

Observation 75b3023e-f670-4324-8b08-aae5b29f9a25 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.146074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.146074Z digest=sha256:e3568b934f5e804de073c34882df94cb4ca13997cae3961cd2070ffcb9a53b20

Observation 26e1751a-a452-4788-af8f-6fa7d246e9df · outbound

This paper cites Dover-lap: A method for com- bining overlap-aware diarization outputs,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Dover-lap: A method for com- bining overlap-aware diarization outputs,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.211149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.211149Z digest=sha256:6f713807362f4bcc231c2a15d5d579e309a119300cc3ded93b7956df0f07eccd

Observation 5d52c1c3-c08a-496c-be86-9f1f01c3556f · outbound

This paper cites The rich transcription 2006 spring meeting recognition evaluation,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The rich transcription 2006 spring meeting recognition evaluation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.172677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:03.345208Z digest=sha256:75d83d20e878d7d266fb00862cbf4ab5b92f2b754b2beb6bf675a79e58a80058

Observation d7c11625-c882-412a-bdda-0e38ad429a5e · outbound

This paper cites Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.921969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:03.405144Z digest=sha256:5f551364ef99d7800bb880e048b11361d267a26506804e91233165b105452784

Observation 0458f9e8-8c9c-4667-956e-6af705fe8edf · outbound

This paper cites GPU-accelerated Guided Source Separation for Meeting Transcription.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition GPU-accelerated Guided Source Separation for Meeting Transcription

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.476909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.476909Z digest=sha256:7d0bab5c4108fb5527804b6b78af2130ed1008692a3e5c3020c4d37564cc4ffa

Observation 92676c42-0012-47ee-9ef8-aa8e151969eb · outbound

This paper cites The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.650135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:03.600303Z digest=sha256:b4c4b518df11bfeb2cd724b8d2759aefc37b01a30e751df7c7e709b3b6c49577

Observation 90723b50-6955-4e15-aa8a-5229bc7d5a0b · outbound

This paper cites VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.767045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.767045Z digest=sha256:240b1e7d539307a4b196780adfeb1c0c3a5a4384c94137c2c7a1606d769015bf

Observation dc3d2ccd-9969-41dc-ac57-55fd0e06b23b · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition VoxCeleb2: Deep Speaker Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.878807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.878807Z digest=sha256:e5421e4e670426fe28034d93c184abad6b0acb44fc4f9bf9a506249264580359

Observation c7e3d097-71db-4f79-bc87-2ea03ec6a50e · outbound

This paper cites Kespeech: An open source speech dataset of mandarin and its eight subdialects,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Kespeech: An open source speech dataset of mandarin and its eight subdialects,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.293656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:03.938904Z digest=sha256:895215a9edddc64f119256eb7dc01cb7e460d02b5be2cb4030aa82197b7227a4

Observation 5ad194ca-b72e-4f4d-9468-e213687d7f26 · outbound

This paper cites 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.029122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.029122Z digest=sha256:ca89272310b890d09b103be16f3c2c4ab87cf94b070d894a464a179f396580c2

Observation 2fbc7da8-9942-4d95-b08a-e464e0796a9c · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.140000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.140000Z digest=sha256:8bdcaa5c5a141f7dc65a50af2a86214b138091b1640529c87cfc56ea2469dde2

Observation f9b28de0-ed4d-4341-9501-1ae12bebade2 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.219246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.219246Z digest=sha256:f2d7f8667a5cf817b8e704164d1303016909c693d85acbbb2637104dcb634d9a

Observation 498eb297-5d96-4a18-9d4f-b1523a25db7b · outbound

This paper cites The ami meeting corpus,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The ami meeting corpus,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.308158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.308158Z digest=sha256:f3b5e1febdc959c0627c12bcc26d9235f82c6312f20bfb32ddfc48068257a994

Observation 49055ca4-5d30-4e6b-b90e-6adb66cce4ce · outbound

This paper cites Lip-reading with densely connected temporal convolutional networks,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Lip-reading with densely connected temporal convolutional networks,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.390574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.390574Z digest=sha256:ad5b134ed94fe68adf5897d506eb1702f104cdd1cc1c120087a5ec4f0ade9010

Observation ca19d67a-28fd-4b16-9564-edaf8e684aaf · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Robust speech recognition via large-scale weak supervision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.485577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.485577Z digest=sha256:3c5b9db9105effd4ecd340c0d5210a80e5f201b5d183c7ac0b43228643463334

Observation 6dbbfdd8-4a91-4c43-be47-b2a52ce12a48 · outbound

This paper cites Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.746106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:04.559893Z digest=sha256:e75cf342df9f609406c64a6143c4330e1afab2429b86868bbe10e0e4265245ba

Observation a70e8cee-0bec-4546-b2c3-3b0c3e71d1fc · outbound

This paper cites Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.678989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.678989Z digest=sha256:b6c426b3c472c87d50727d72bdf0920089daa7aea0d885705c414f175278dccc

Observation 909b7d81-34cb-44bd-b0dd-6df558509c42 · outbound

This paper cites Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.846296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.846296Z digest=sha256:cbdabd8d8fa182746dbeceb44d687b1e1a3de5efb4bdfcafb89d36ab656afa79

Observation 1494347b-02b1-43d7-9318-d5ff6fcb3771 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.928529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.928529Z digest=sha256:1213c49cac8ae0a6fc92ec14e1f00334e47c80012a9a2722523fe72332cde2ff

Observation fc4ac28f-13cd-4e27-8b80-0d5abb96a892 · outbound

This paper cites Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.514736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:05.021485Z digest=sha256:7c74965875e587842cb22442aa40be91fc078fbfa86b6effd4661c89de5fe5a3

Pith citing papers

Observation d516a4d6-1b5c-43fa-bd6f-344aae0d6d84 · inbound

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition cites this paper.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:00.946933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:00.946933Z digest=sha256:d028ef62bc787f23aac46c1fd7feafa65dfbc5e248040002f9019307a2bc2c83

Observation 02c18a0e-adb3-4d68-9574-028150db69db · inbound

Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge cites this paper.

Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:26.572189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:26.572189Z digest=sha256:787f1767e691bfa72e1d9af51d5d663d54dd30b2c6c5c0da3ca693d2e1494c73

Observation dca6c443-3407-4b0b-92fb-977eb884cbde · inbound

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset cites this paper.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.777201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.777201Z digest=sha256:623398c14538e0ead60c50f767bcbfb2c05db6a2ab0e7b721c041fed0db25911

Observation 06774b7d-55b1-42a7-abae-e6f845a405cf · inbound

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs cites this paper.

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:09:51.398870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T08:06:45.477344Z digest=sha256:c58043414ca9e54dfff0aa415f9cfaa1b41f022a08acc461ee92c2775afd7384

Observation 038ca04a-f233-43de-a3f7-8b08fa5f8d8a · inbound

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models cites this paper.

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:21:12.615570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T09:18:53.285951Z digest=sha256:5c3837053a47ff6ad9f43b30e0fb32730f47e5654565e6630389df7ebbe27e71

Observation 38be85af-3c4d-473d-960a-35da92a73616 · inbound

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings cites this paper.

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T13:06:07.806541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:06:07.806541Z digest=sha256:681a3989724e38ad957ebf1d6ad0a98680ed566e80e9becac77d41a7331f33cd