Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:05.021485Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2505.13971.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:05.021485Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:00.946933Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T08:09:51.394586Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d516a4d6-1b5c-43fa-bd6f-344aae0d6d84 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d182b9a8-a55d-4e53-a58c-b0807a475f2c · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4617894b-5ab6-4b5d-b4c4-abceef6f4b8f · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition who spoke when
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 63d2ca25-9035-4f8f-9644-6e043c78bd31 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af8c5e2b-7ce1-461d-9630-7081a33c5d9e · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition (2), whereNspk is the total number of speakers in the session
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fd03b07a-e841-49f1-965a-560307c032da · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fabb067c-e0b7-4de2-9160-4bc51668a496 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 095c7367-1392-4e02-a4bd-1df23564b5f2 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ee43906-d77d-4382-a9c1-11056bb6132d · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94b03a91-ce32-48d7-8791-45fde30cb4a5 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 474dda51-9528-4bbb-86e4-476a84c651c5 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24dda821-8acf-4170-af19-ffb35fec1c65 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Continuous speech separation: Dataset and analysis,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9c3e128-15c4-4b6e-ac08-f377e39382e2 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Lib- rispeech: an asr corpus based on public domain audio books,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef7adc73-dd1c-4ab8-b3d1-cf8b99dd345f · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b477c2d-8e4a-450f-bd4f-280b2da4535e · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bea85d7f-dd75-41a5-bfda-ee982ef8a884 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c480d17-e3da-41ee-96c7-d54cb0103dbc · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4e4131e-df16-4fd3-abab-13d0adc854ab · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75b325c-4bf8-4d38-a95a-702aac5698b2 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition End-to-end audio-visual neural speaker diarization,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 75b3023e-f670-4324-8b08-aae5b29f9a25 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Conformer: Convolution-augmented Transformer for Speech Recognition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e1751a-a452-4788-af8f-6fa7d246e9df · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Dover-lap: A method for com- bining overlap-aware diarization outputs,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d52c1c3-c08a-496c-be86-9f1f01c3556f · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The rich transcription 2006 spring meeting recognition evaluation,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7c11625-c882-412a-bdda-0e38ad429a5e · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0458f9e8-8c9c-4667-956e-6af705fe8edf · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition GPU-accelerated Guided Source Separation for Meeting Transcription
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92676c42-0012-47ee-9ef8-aa8e151969eb · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 90723b50-6955-4e15-aa8a-5229bc7d5a0b · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc3d2ccd-9969-41dc-ac57-55fd0e06b23b · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition VoxCeleb2: Deep Speaker Recognition
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e3d097-71db-4f79-bc87-2ea03ec6a50e · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Kespeech: An open source speech dataset of mandarin and its eight subdialects,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ad194ca-b72e-4f4d-9468-e213687d7f26 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fbc7da8-9942-4d95-b08a-e464e0796a9c · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b28de0-ed4d-4341-9501-1ae12bebade2 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 498eb297-5d96-4a18-9d4f-b1523a25db7b · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The ami meeting corpus,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49055ca4-5d30-4e6b-b90e-6adb66cce4ce · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Lip-reading with densely connected temporal convolutional networks,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca19d67a-28fd-4b16-9564-edaf8e684aaf · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Robust speech recognition via large-scale weak supervision,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dbbfdd8-4a91-4c43-be47-b2a52ce12a48 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a70e8cee-0bec-4546-b2c3-3b0c3e71d1fc · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909b7d81-34cb-44bd-b0dd-6df558509c42 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1494347b-02b1-43d7-9318-d5ff6fcb3771 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4ac28f-13cd-4e27-8b80-0d5abb96a892 · outbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d516a4d6-1b5c-43fa-bd6f-344aae0d6d84 · inbound
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c18a0e-adb3-4d68-9574-028150db69db · inbound
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca6c443-3407-4b0b-92fb-977eb884cbde · inbound
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06774b7d-55b1-42a7-abae-e6f845a405cf · inbound
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 038ca04a-f233-43de-a3f7-8b08fa5f8d8a · inbound
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38be85af-3c4d-473d-960a-35da92a73616 · inbound
MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.