Pith. sign in

Paper Citation Record · LEDGER

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2505.13971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13971 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:05.021485Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:00.946933Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T08:09:51.394586Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d516a4d6-1b5c-43fa-bd6f-344aae0d6d84 · outbound

This paper cites The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:00.946933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:00.946933Z digest=sha256:f3e6fd6f7a31648d8173fe6223a6d79b688cb2d6ded3af955f0c5f783335f02d

Observation d182b9a8-a55d-4e53-a58c-b0807a475f2c · outbound

This paper cites Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:10.405861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.015198Z digest=sha256:c86d0d1053be2cb8c99511b863de8a09cd3fcc16e6748b7e81adcceb374674ae

Observation 4617894b-5ab6-4b5d-b4c4-abceef6f4b8f · outbound

This paper cites who spoke when.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition who spoke when

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:10.171540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.131980Z digest=sha256:f1f22eba6d470b3d64f70e27320adf3487704e21d5619d74973b1e72653e24ff

Observation 63d2ca25-9035-4f8f-9644-6e043c78bd31 · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:10.001010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.253337Z digest=sha256:a9e726e791c23031c27aee72ecf78259a867a003a0e4c20d9833f7954f85b4e5

Observation af8c5e2b-7ce1-461d-9630-7081a33c5d9e · outbound

This paper cites (2), whereNspk is the total number of speakers in the session.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition (2), whereNspk is the total number of speakers in the session

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.796298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.356018Z digest=sha256:594f6f283807177fa94f2d53601a274faefe1c47264abbad521b31a5f09d03fe

Observation fd03b07a-e841-49f1-965a-560307c032da · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:09.545159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.496388Z digest=sha256:3dccbef0fa4187fce5c9141eee84ec5c7b80a4b2758cc8b1d9e2ded4ca246c89

Observation fabb067c-e0b7-4de2-9160-4bc51668a496 · outbound

This paper cites A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.281598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.616838Z digest=sha256:eee4a66beda0b9eeedbceade6a97e27fb31372929031bf2183e06ca196f0c694

Observation 095c7367-1392-4e02-a4bd-1df23564b5f2 · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:09.111934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.738485Z digest=sha256:1b7273a56252f0f45361a4d08f98b1cb75e83809f50072560f44a6a3d561cf54

Observation 1ee43906-d77d-4382-a9c1-11056bb6132d · outbound

This paper cites The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.828947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.825645Z digest=sha256:8ad3a74bbfcff61bfe6c426125484502b5a16915b8406e7a9dbdb90f9cf329a8

Observation 94b03a91-ce32-48d7-8791-45fde30cb4a5 · outbound

This paper cites M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.634867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:01.970675Z digest=sha256:91772d4303efba7314e70ac4813c4a03be1061a0643f1a24a74f13ef5cc48727

Observation 474dda51-9528-4bbb-86e4-476a84c651c5 · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.079339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.079339Z digest=sha256:83609c6288b65392aa508ec4ee394b663d33f25f2ee6c241d726e163717ac09c

Observation 24dda821-8acf-4170-af19-ffb35fec1c65 · outbound

This paper cites Continuous speech separation: Dataset and analysis,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Continuous speech separation: Dataset and analysis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.436561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:02.211446Z digest=sha256:be04c4b1baa71bb612e843447e52c34ebc3f56d785a7e52156a97c60f21d035c

Observation e9c3e128-15c4-4b6e-ac08-f377e39382e2 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Lib- rispeech: an asr corpus based on public domain audio books,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.305820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.305820Z digest=sha256:7b9a8c811da2100f55bd42220b2a06f18afed929fdf556475e6ab05d3ea695df

Observation ef7adc73-dd1c-4ab8-b3d1-cf8b99dd345f · outbound

This paper cites CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.186566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:02.459989Z digest=sha256:cf15ac2af6f2f770d755caa6bfb450bb0ba1c48127eab8291b52b6a3806d17fd

Observation 5b477c2d-8e4a-450f-bd4f-280b2da4535e · outbound

This paper cites The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.989009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:02.614902Z digest=sha256:468f1aeed41621f78e404b25360aa0aa0189b7bda45b78990f79ec435e40db5c

Observation bea85d7f-dd75-41a5-bfda-ee982ef8a884 · outbound

This paper cites The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.797124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.797124Z digest=sha256:822efed051ffacf2621caccda3e07f72f4cfb928564f0517a24b0bb8563c0a3f

Observation 9c480d17-e3da-41ee-96c7-d54cb0103dbc · outbound

This paper cites The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.764255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:02.876803Z digest=sha256:1652502d809018526b684f1d2d20c8d0e6c41cc6641b3514326d1d98666c7ba6

Observation c4e4131e-df16-4fd3-abab-13d0adc854ab · outbound

This paper cites MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.014747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.014747Z digest=sha256:54d37ed6b1f5ac078c5158f6076aad6a0ed9564c9b2312e445a82a92e26555ba

Observation b75b325c-4bf8-4d38-a95a-702aac5698b2 · outbound

This paper cites End-to-end audio-visual neural speaker diarization,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition End-to-end audio-visual neural speaker diarization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.550360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:03.076635Z digest=sha256:fd5a42c013e3f973f22baa24eba4edc141bd9deffe167109863ff995d794bf63

Observation 75b3023e-f670-4324-8b08-aae5b29f9a25 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.146074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.146074Z digest=sha256:a695b366c0bb22c12dd359b5ace6cdb5d44da82fd7c84a6d3358e349d4085d7b

Observation 26e1751a-a452-4788-af8f-6fa7d246e9df · outbound

This paper cites Dover-lap: A method for com- bining overlap-aware diarization outputs,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Dover-lap: A method for com- bining overlap-aware diarization outputs,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.211149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.211149Z digest=sha256:92545ff95cbfb4173b19732b3a4cc537dc9a09ffda42b972f543492d8c24c84a

Observation 5d52c1c3-c08a-496c-be86-9f1f01c3556f · outbound

This paper cites The rich transcription 2006 spring meeting recognition evaluation,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The rich transcription 2006 spring meeting recognition evaluation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.172677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:03.345208Z digest=sha256:950ed1b3defc916d411e16f3af36da28c00152ddcda2a0382299dd750aba1ad6

Observation d7c11625-c882-412a-bdda-0e38ad429a5e · outbound

This paper cites Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.921969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:03.405144Z digest=sha256:c0c893584ac3190589628822665f93694d26f122c81cae49c488623452897298

Observation 0458f9e8-8c9c-4667-956e-6af705fe8edf · outbound

This paper cites GPU-accelerated Guided Source Separation for Meeting Transcription.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition GPU-accelerated Guided Source Separation for Meeting Transcription

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.476909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.476909Z digest=sha256:1ed08f8b4a19dcda14f5e3ee5a62b5bb56485f18b2c67ba799cbbff739a314e5

Observation 92676c42-0012-47ee-9ef8-aa8e151969eb · outbound

This paper cites The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.650135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:03.600303Z digest=sha256:e89e5a916e1214d7bb924c355ccc2382dd2654d5afbc08483d8cdd3c1e84f140

Observation 90723b50-6955-4e15-aa8a-5229bc7d5a0b · outbound

This paper cites VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.767045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.767045Z digest=sha256:f95fdd6b30a1987e0f8bf30bc8d480ba0b11dfc952146216a82198c2f8cb29cb

Observation dc3d2ccd-9969-41dc-ac57-55fd0e06b23b · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition VoxCeleb2: Deep Speaker Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.878807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.878807Z digest=sha256:4c2a8a25b46c538147aae1676d40986a641d56bb46c67ef4daf78722f9e52f2f

Observation c7e3d097-71db-4f79-bc87-2ea03ec6a50e · outbound

This paper cites Kespeech: An open source speech dataset of mandarin and its eight subdialects,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Kespeech: An open source speech dataset of mandarin and its eight subdialects,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.293656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:03.938904Z digest=sha256:bf94303195b24e22d5f7374dccf6e85819d2b567e172ed04f11484d5556b138b

Observation 5ad194ca-b72e-4f4d-9468-e213687d7f26 · outbound

This paper cites 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.029122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.029122Z digest=sha256:58d86a030faa85aba4f359bc8567199a56a45bfdcb38aef284e796f46f04531c

Observation 2fbc7da8-9942-4d95-b08a-e464e0796a9c · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.140000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.140000Z digest=sha256:9471df6c1bb3c6e01d0ebf32e907dcadd54fc554759919712ed6191ace2c0e7b

Observation f9b28de0-ed4d-4341-9501-1ae12bebade2 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.219246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.219246Z digest=sha256:8a2ef36ee0c8795cec332c79ed2c204d996d850c728aea5e085b452c101ccb5e

Observation 498eb297-5d96-4a18-9d4f-b1523a25db7b · outbound

This paper cites The ami meeting corpus,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The ami meeting corpus,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.308158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.308158Z digest=sha256:dc02a384e2ff47eac133fc26f776a928ebc869455d7dbb869af37ffb9c704c01

Observation 49055ca4-5d30-4e6b-b90e-6adb66cce4ce · outbound

This paper cites Lip-reading with densely connected temporal convolutional networks,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Lip-reading with densely connected temporal convolutional networks,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.390574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.390574Z digest=sha256:0e31e3c4468d7cd979e4702d711e35db2edcb739820a8ea0583001b3ec98de1e

Observation ca19d67a-28fd-4b16-9564-edaf8e684aaf · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Robust speech recognition via large-scale weak supervision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.485577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.485577Z digest=sha256:2eb72d2f46d74847d6efe488f56c20eef889e63bb1d9cb3081968a69cedb1dcc

Observation 6dbbfdd8-4a91-4c43-be47-b2a52ce12a48 · outbound

This paper cites Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.746106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:04.559893Z digest=sha256:504f734a9eb41e7fbc41f9b9eb51fb02eb14358d2f20b963699f839ce42d92a9

Observation a70e8cee-0bec-4546-b2c3-3b0c3e71d1fc · outbound

This paper cites Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.678989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.678989Z digest=sha256:c8b129d58b7576934f1e0aedb45fc61909e8881ad4b3baafd1c5dbc4246e3404

Observation 909b7d81-34cb-44bd-b0dd-6df558509c42 · outbound

This paper cites Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.846296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.846296Z digest=sha256:35fefcdfafcd9d44fa628a0efcd9b16c30ad70a74f169c401b8f6090c32bbf19

Observation 1494347b-02b1-43d7-9318-d5ff6fcb3771 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.928529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.928529Z digest=sha256:1b424b26c01bc5992946993608417a4d97f88613393bff0956a134c1b8b8f75b

Observation fc4ac28f-13cd-4e27-8b80-0d5abb96a892 · outbound

This paper cites Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.514736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:43:05.021485Z digest=sha256:9259401bce4d184813374056a90c836e005c19d324eb1bf1c0593cf5ff03b5fc

Pith citing papers

Observation d516a4d6-1b5c-43fa-bd6f-344aae0d6d84 · inbound

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition cites this paper.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:00.946933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:00.946933Z digest=sha256:f3e6fd6f7a31648d8173fe6223a6d79b688cb2d6ded3af955f0c5f783335f02d

Observation 02c18a0e-adb3-4d68-9574-028150db69db · inbound

Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge cites this paper.

Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:26.572189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:26.572189Z digest=sha256:32c84804beed8e0d417f83bbf659abf3680e09d7c5852c8aad63ff773bc2bdcd

Observation dca6c443-3407-4b0b-92fb-977eb884cbde · inbound

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset cites this paper.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.777201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.777201Z digest=sha256:329d269e7d5b38f3805c7852c11fda2f3e2f42724943121ade75142fa723130c

Observation 06774b7d-55b1-42a7-abae-e6f845a405cf · inbound

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs cites this paper.

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:09:51.398870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T08:06:45.477344Z digest=sha256:238aeda6fc4d2a0b7dcf0e46faab5bc08aa05e4b4a44c325a42da76101cf0dbe

Observation 038ca04a-f233-43de-a3f7-8b08fa5f8d8a · inbound

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models cites this paper.

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:21:12.615570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T09:18:53.285951Z digest=sha256:82b73faef64d9b2c2c3ac8bd12a84a7cdab89f6da64744c46d6bd441e0ba8a07

Observation 38be85af-3c4d-473d-960a-35da92a73616 · inbound

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings cites this paper.

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T13:06:07.806541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:06:07.806541Z digest=sha256:264b86d8a96b2635ec3f7115a9d9fd2618c2e492f1ce292691616d49030e1987