Pith. sign in

Paper Citation Record · LEDGER

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

As of 19 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 13 inbound Pith citation observations for arXiv:2501.13306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13306 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:21:50.792032Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:33:10.833843Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:49:50.788200Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4845ecc-ab65-4973-bb13-35ee447b6f80 · outbound

This paper cites Data Products , 2024.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data Products , 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.254183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.641410Z digest=sha256:157da0c63f509cab2bef5ac8b73c99c4a20abdc573315d0c8fa150a16e20c103

Observation 0a0579ea-5f4d-4fa8-a1e7-9c9f4ee76847 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.645939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.645939Z digest=sha256:abace3e9d2a9165513c3cbff9bf5b44eb47f574d7c48cf3a38c4576715852dd3

Observation 26405e1f-01c6-4177-9c8e-34b8d827b0a4 · outbound

This paper cites Qwen Technical Report.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.650773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.650773Z digest=sha256:1ebb027d5fba4b1a1c034b02ddac218aeba158b75d2fdef49d629a501b656f44

Observation 1d2f7d2d-bf49-4319-9d59-0b74ce33614c · outbound

This paper cites AISHELL-1 : An open-source mandarin speech corpus and a speech recognition baseline.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia AISHELL-1 : An open-source mandarin speech corpus and a speech recognition baseline

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.241788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.655249Z digest=sha256:9930c7dcbdca1714428c2fd5e96292b29ecbfeaf4712ac0a77f284dceab46805

Observation 46728167-1fc3-45b8-b624-2198cb239ab6 · outbound

This paper cites IEMOCAP : Interactive emotional dyadic motion capture database.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia IEMOCAP : Interactive emotional dyadic motion capture database

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.229299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.660136Z digest=sha256:abb91d8667056f8f88635a08e646bff8cda60f5556ccec940253c29a419698fe

Observation 52440562-690e-4e16-84cb-b880bc7a35ee · outbound

This paper cites MSP-IMPROV : An acted corpus of dyadic interactions to study emotion perception.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MSP-IMPROV : An acted corpus of dyadic interactions to study emotion perception

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.217573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.664124Z digest=sha256:1688aea4e73f56ddec513357a73f02356ca64d3bbfef576688cea146a846a9fb

Observation 15110fa2-db69-4e3d-a69b-0a745922a32d · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.668642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.668642Z digest=sha256:06e71e9a94b8b88d535385a6e0308f071412f826b080e9bd58695c7f799c8506

Observation 953d4eb8-70f3-4527-b300-24d2dec5e61f · outbound

This paper cites Qwen2-Audio Technical Report.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen2-Audio Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.672430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.672430Z digest=sha256:c644d11ed5987b063e7dddf79b56d8044b57ead14488fcac804edc9b8c211718

Observation 2b3e2324-61fc-46a6-8ea9-9c3b754f2a8b · outbound

This paper cites Data products, 2024.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data products, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.205841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.676243Z digest=sha256:38a5aeaac6733d806397d5d835df81d3bb525275a6b651c399268c6e7d85edb5

Observation 0013ddfe-9b79-4433-88e9-fedde57b0f45 · outbound

This paper cites Data products, 2024.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data products, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.194923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.679557Z digest=sha256:8ffd72d05cac309df162e3975a6cfc9d21d6a03301ec2713e87d19fa2e7ffc6e

Observation ef476ca6-1f6f-4dad-8424-4384230db56e · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.682871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.682871Z digest=sha256:edaa7bea727e11bdb58b8b52f715dc6e55e0fea281200b97d5af10b844b812fe

Observation 02eea089-deb6-4bff-ab24-2e277fa8f5e9 · outbound

This paper cites Gemmeke, Daniel P.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Gemmeke, Daniel P

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.184074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.688136Z digest=sha256:6340b7f4d06956d4ee7cd43e3d330db1412db71bbfd9fe5cdcc21f7b6c19f80a

Observation b763e6cd-8fb7-45a2-ac99-ef753df47ed6 · outbound

This paper cites Vocalsound: A dataset for improving human vocal sounds recognition.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Vocalsound: A dataset for improving human vocal sounds recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.173901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.692800Z digest=sha256:4dfe7fc29fcae7f57fb2bb4d8c925786687a13454a626a991d509e07054f0d9f

Observation f7be6a17-c051-4edf-a4c8-4012dc488c1a · outbound

This paper cites LoRA : Low -rank adaptation of large language models.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia LoRA : Low -rank adaptation of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.163633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.697178Z digest=sha256:660987c81dd4bfa98b5c685c70c4edcad0806924baee8ae7bd2f91cff7afed98

Observation d2030894-3781-4c39-91b1-699135cd2541 · outbound

This paper cites Datasets, 2017.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Datasets, 2017

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.153105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.701458Z digest=sha256:f8fc2511bbb3b3ddd49faebfaab1c3e1378970ae35b1dca76b263204cbd509d5

Observation b228b375-8c03-4806-9680-93ae9f01d650 · outbound

This paper cites Schuller, and Jianhua Tao.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Schuller, and Jianhua Tao

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.142135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.705820Z digest=sha256:524b5906414bf258088d0cac946a3e82aace2ffe0f260ff283e1817ec71996db

Observation f66d6d84-5633-41ca-9be7-69b3c0021db3 · outbound

This paper cites Emotion2vec: Self-supervised pre-training for speech emotion representation.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.131662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.709957Z digest=sha256:c2886d03f782b8016b2a3791236fa33aca080c53023b2b8acf53e878152f1261

Observation fdfd17ed-1cf7-4c0e-be38-c59aaeadb688 · outbound

This paper cites The MSP -conversation corpus.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia The MSP -conversation corpus

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.120340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.714460Z digest=sha256:caf0bf59636bb51298d76e2582a546ea3216c7b7e387fc77f04da998d5ba253e

Observation f9167552-5a2f-4ed4-9def-0c061fd591f3 · outbound

This paper cites MAGICDATA mandarin Chinese read speech corpus, 2019.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MAGICDATA mandarin Chinese read speech corpus, 2019

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.107988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.718846Z digest=sha256:c92b2b2cff2511e86dcbda38c5e2dacf31f1931d96345add8a5786bf8804e08b

Observation a9ba5083-904c-4640-9995-ae09ad4aef9e · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Librispeech: An ASR corpus based on public domain audio books

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.096486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.724151Z digest=sha256:4847c3b8fd4dbd2a6c79ffb9a8ed2b4f03f7043707aba4a77bc8f3afac90585e

Observation 40c2663a-8e81-4e98-b77a-ac501ac14610 · outbound

This paper cites Reproducing whisper-style training using an open-source toolkit and publicly available data.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Reproducing whisper-style training using an open-source toolkit and publicly available data

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.084881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.728388Z digest=sha256:48b88c0263f26b929ae1eefebf3c54c349cb5778e8fa8cd43bbc8aaf5ff3f023

Observation a5df2899-512f-413b-933e-4b6cfa6d6f51 · outbound

This paper cites an unresolved cited work.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:21:51.072413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.731940Z digest=sha256:42e36940811c06e47767d145fd342fe62537d5bec5d6783b1c72cf76920e823f

Observation 0c1bfbc9-4cb3-435b-9387-85f3af8b9463 · outbound

This paper cites MELD : A multimodal multi-party dataset for emotion recognition in conversations.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MELD : A multimodal multi-party dataset for emotion recognition in conversations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.061704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.735946Z digest=sha256:0e396a92a3205e97fcd0208efc2c02ca7ea6186d2f1786e4656f96d12fdc3af1

Observation c4678dbb-125c-49d6-8e58-cd612f3a968a · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Robust speech recognition via large-scale weak supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.051351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.739872Z digest=sha256:6fde23f78a81f7fa05f6fcb6532d4ddb4bf2833e07ca5bafe500c956f41678e9

Observation 495dd3a5-d28c-4dd6-a54d-f0e27ab3d062 · outbound

This paper cites Nonspeech7k dataset: Classification and analysis of human non-speech sound.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Nonspeech7k dataset: Classification and analysis of human non-speech sound

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.039921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.743915Z digest=sha256:a77ab091416d1df0abda385f488e2a15264adeb08b47c6aa7e5a5314eda7934e

Observation 64e3ad61-25f5-4207-90f2-ffa3e92de62b · outbound

This paper cites The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.748041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.748041Z digest=sha256:2c267fa5249beeb1ee302401ff72287928fffd7f1b1dda2d33eab681d9ea1740

Observation e20f0866-7a74-4ce1-8c57-5d36fc155563 · outbound

This paper cites Achieving timestamp prediction while recognizing with non-autoregressive end-to-end ASR model.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Achieving timestamp prediction while recognizing with non-autoregressive end-to-end ASR model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.026269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.751915Z digest=sha256:b38ff11af951e4ad365d276f10b413f872f33e5eedf82241cff13a388b2c3643

Observation afe57756-a2d3-470d-80f4-fd23a3dffd5b · outbound

This paper cites TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:21:50.833353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.755899Z digest=sha256:60cdf2c6038e44a6eec4660fe443d8f6270e00c91ad32b065ef7606b5c2f3a61

Observation 84ec1fc9-4e59-48ac-8f98-fc5b362f883e · outbound

This paper cites PandaGPT : One model to instruction-follow them all.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia PandaGPT : One model to instruction-follow them all

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.014485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.760638Z digest=sha256:191ff9c6da89084100e27defbe499213ed1672285681bbef67e6fedc2632d045

Observation 5fcabe55-126a-48c5-bc0c-739cc90a3f2d · outbound

This paper cites SALMONN : Towards generic hearing abilities for large language models.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia SALMONN : Towards generic hearing abilities for large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.002925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.764260Z digest=sha256:908be733a4d2f4719f8d4e786dba89d1923a52411cad395b7bccfff42dbb7577

Observation 2f9dd831-0c14-453c-9f29-ec7c9ad4fe32 · outbound

This paper cites Kespeech: An open source speech dataset of Mandarin and its eight subdialects.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Kespeech: An open source speech dataset of Mandarin and its eight subdialects

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.992013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.767988Z digest=sha256:f44d4bb1113e380d7c83332940b327bcf438e152f4282fa15b45e355134caf12

Observation 5eef7f39-a106-4968-9c33-982a1ac3b2c2 · outbound

This paper cites Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.980216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.771903Z digest=sha256:f2184a30423cd6d13daa5f8dcf994f72d8044709a1d88d5977ae8fd8a24ff6bf

Observation 52d13f61-9502-479a-936d-b5b890a9a7ae · outbound

This paper cites Attention is all you need.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Attention is all you need

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.967826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.775557Z digest=sha256:beda6bb1ee1998552386662019f8d191943750d6a07c4353b399d3bd6c974b45

Observation caff80b2-1551-4bc4-b5e7-265f7e582f6b · outbound

This paper cites A large-scale Chinese short-text conversation dataset.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia A large-scale Chinese short-text conversation dataset

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.955944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.779084Z digest=sha256:36659315dc871efac30f3c4efce52739aae26f78aa8089c23410d5f27bd64e97

Observation 75834008-8b63-4873-963e-aa9b98735c35 · outbound

This paper cites WENETSPEECH : A 10000+ hours multi-domain Mandarin corpus for speech recognition.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia WENETSPEECH : A 10000+ hours multi-domain Mandarin corpus for speech recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.943291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.783589Z digest=sha256:f34dca72381f3d5bba4437a26282eee7b858f51642b2c2bbc035b45a2b3776ec

Observation 98512b3b-063f-4236-be2d-d25226d42cf0 · outbound

This paper cites M 3 ED : Multi -modal multi-scene multi-label emotional dialogue database.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia M 3 ED : Multi -modal multi-scene multi-label emotional dialogue database

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.930655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.788155Z digest=sha256:0d63af93a0897fe712007cef42d4aa6c77bea2daae47b0ac68275e3676b60540

Observation 26a6b988-ecf9-42c2-8381-48f36e2c8c64 · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.916362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.792032Z digest=sha256:278a5d01d1f0c8a71ef3c0b5d12cf75a443e6d8f336972b3782e2c301af49070

Pith citing papers

Observation 06cb53d3-6404-4a88-af11-6ebc594f1ca1 · inbound

MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind cites this paper.

MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:33:10.833843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:33:10.833843Z digest=sha256:93c055c056c0c4c0061257f5de384b3e53cc2f8407735e9d4d8f597923c3db8d

Observation 203b544f-3836-4423-a139-4e5a5f95024d · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.283015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:f70a0c815496d9ff1d6785bf7ec32180faa6f999f51a227a8fa0ff4d4eb9c3ea

Observation 314f9479-bc79-4f09-bcb5-5fa240649b68 · inbound

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models cites this paper.

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.939607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.939607Z digest=sha256:992daf777e49c8c7337a4824601f6b0f5874fe0b58b9273d5d9d0981877b009a

Observation 27882c48-9fd5-4dce-80ee-c7519ae63114 · inbound

BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation cites this paper.

BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T20:59:46.041859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:59:46.041859Z digest=sha256:014496b3aec2d774626aaf6ef47b8c30f2488d15ff52ee6bdbf8da838e1b7c34

Observation 787c8920-5478-449b-bb67-7791784322b4 · inbound

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue cites this paper.

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T21:03:09.640862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:03:09.640862Z digest=sha256:d99d671d20b30db2bbc8d647020cefc77deda64b3f555f14285506e45fc00cbd

Observation 93284ac9-394a-42dc-9ced-1d4b75bc47dd · inbound

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages cites this paper.

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.228057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T16:27:37.596817Z digest=sha256:cd7fd7ff2ff56b07a4aec67ce8dd896c2a80b1bac6dafce4a939d041e71bafbf

Observation 267ba0df-de4f-406e-956d-5b4178a6c813 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:38.479016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:38.479016Z digest=sha256:b74405f63ef322d7f7a56c1aef0f0be3befbbbbbb7aaedf90b9c04c37d675484

Observation 196a7acf-ff2b-4f85-a4b7-7f2990d536e4 · inbound

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models cites this paper.

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.898403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:24:27.118694Z digest=sha256:285b9252377e2e1a6a6bc49257b69b996b07772b792a5226d9bae9803cc2bfcb

Observation 6ead0d2e-a824-4f9f-951d-49a72e4765b3 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.342560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:7eba07ec9ab877fe9ad818f06223618c3ba70c9a02b36803b26618eec9b492ee

Observation d2db76cc-80e9-4dfd-93d8-fbf512f19a12 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:1a2ebb96cc806ad2e00da7590a23dacd2b5ddde9b8d5198d0208ddd6f85c4fd9

Observation ae99d2f1-9b81-426c-9ce1-96113b53d4bc · inbound

Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages cites this paper.

Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:31.006517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T04:25:37.725216Z digest=sha256:a8f56c7a98f4af01c64c8aff5baee76b52dcda375d3e816b19ff589efc147776

Observation 760bc7e5-e472-4085-b43d-a2c487a220c8 · inbound

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios cites this paper.

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:49:50.790654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T07:36:34.307652Z digest=sha256:4628f7314a2d0be8a3468928e95da895fd4e95320ec514a4660d8dfa7f122d2e

Observation 459daaf3-3a19-4091-a8cc-2608f1eb74b1 · inbound

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond cites this paper.

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T05:50:29.618737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:50:29.618737Z digest=sha256:4c51eef806c1002bfd1d7c0c60c5809711c2d665b83d6641d914454400641836