Pith. sign in

Paper Citation Record · LEDGER

MOSS-Audio Technical Report

As of 5 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 6 inbound Pith citation observations for arXiv:2606.01802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01802 v3

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T13:05:29.813707Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:35:56.402445Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T14:53:55.822643Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact33
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af73f21c-68b5-4ae8-9f05-e1997837d6fb · outbound

This paper cites MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark.

MOSS-Audio Technical Report MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.096484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:557a1db0faf11e22db83dce65518d6883354fea19fa6c439da6924ba51da602e

Observation f448aabb-33d2-4c3d-9271-026bb43374f2 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

MOSS-Audio Technical Report Audio set: An ontology and human-labeled dataset for audio events

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:e653bbe0ef98f7b37f82497199eb3a506d55abe0d63bca6df385e9d66893d11a

Observation ae8c38fa-9e02-4ccc-9d8c-db4a3f5fe9b0 · outbound

This paper cites A udio C aps: Generating captions for audios in the wild.

MOSS-Audio Technical Report A udio C aps: Generating captions for audios in the wild

Reference 3

Resolution
verified exact
doi, observed 2026-06-28T13:12:13.681711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:a52ebec92080291fcfb1d77953092816adf6fd9334076f1c097fcd69e61b4b2b

Observation 42c08a71-b4c5-4add-817f-8cc304eda715 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing 29, 1161–1173 (2021).https://doi.org/10.1109/TASLP.

MOSS-Audio Technical Report IEEE/ACM Transactions on Audio, Speech, and Language Processing 29, 1161–1173 (2021).https://doi.org/10.1109/TASLP

Reference 4

Resolution
malformed identifier
doi_truncated, observed 2026-06-28T13:12:13.686139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:f5a3846bdf5ba69f4f855976d3f8d409f1866dd5186d72dc0f32353dd4a706e4

Observation 68aa2f45-0b4c-4c0e-aaa6-a43f8b2c6865 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

MOSS-Audio Technical Report Robust speech recognition via large-scale weak supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:12150c8927bd5cc882b7503384ad3f0bcb8e14ddedab72eb497a3745af8b9331

Observation f4ca3108-7a9c-4749-b2ad-8a07fd214955 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

MOSS-Audio Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.088047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:74fc48f23a1265055e363294d5abcd82cc065f3e1ff40610f479dfa290446286

Observation 9af80082-ed7b-4cdd-beda-7d9c93e5f27a · outbound

This paper cites SALMONN: T owards generic hearing abilities for large language models.

MOSS-Audio Technical Report SALMONN: T owards generic hearing abilities for large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:536a24fb6e58882963f7b2e63dcb821d129ccfdd22536dfc19c3b8f21095baeb

Observation ed21e055-1ae8-4962-8023-f7d1eb8b0350 · outbound

This paper cites Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities.

MOSS-Audio Technical Report Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:c552f6885e6a42f28e066a7502717eafa3b5d7599b4a6fe75acd47e6a66c40ed

Observation 3dadd2c6-621d-4838-83a5-e8afc13519c4 · outbound

This paper cites Qwen2-Audio Technical Report.

MOSS-Audio Technical Report Qwen2-Audio Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.094371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:edb759f5dfa9e9af44c5f15d816c1cfa6cd9458536751070262da6f60a250065

Observation 80a6f3c8-f54f-46e1-b2fc-7ba9f9949b4e · outbound

This paper cites Qwen2.5-Omni Technical Report.

MOSS-Audio Technical Report Qwen2.5-Omni Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.117896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:8c4524fcfd7c67d032d2e9f4db1cd5f9a6900da20a1781713c42c4fd3bcfb7fb

Observation fa56ac27-c8be-475c-b64d-18ff7d03ab88 · outbound

This paper cites Sakshi, Oriol Nieto, Ra- mani Duraiswami, and Dinesh Manocha.

MOSS-Audio Technical Report Sakshi, Oriol Nieto, Ra- mani Duraiswami, and Dinesh Manocha

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:556b6c69f4b409e08826a44da6445b2c2d30870d1d6c35c6a5b0240f4c0ce9bd

Observation 0c0ed9ed-f01f-4924-a57e-cdbe6b4f47c5 · outbound

This paper cites Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music.

MOSS-Audio Technical Report Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.085342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:778b93a51efdf841291ec1a6246aaf24bd783d0ec6ef3ee23d8e2edf9c230883

Observation cbf4515f-2c82-412e-82a8-fcc8532e758e · outbound

This paper cites Enhancing temporal understanding in audio question answer- ing for large audio language models.

MOSS-Audio Technical Report Enhancing temporal understanding in audio question answer- ing for large audio language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:26b38b88d579375a7845be81275d9422b0945ced5b8862f9b0b9491df0235c51

Observation 39c3e04d-08df-463a-af71-f8221bc326d2 · outbound

This paper cites Sakshi, Utkarsh T yagi, Sonal Kumar, Ashish Seth, Ramaneswaran Selvakumar, Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, and Dinesh Manocha.

MOSS-Audio Technical Report Sakshi, Utkarsh T yagi, Sonal Kumar, Ashish Seth, Ramaneswaran Selvakumar, Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, and Dinesh Manocha

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:4d174998dca00c36c6b947217480199e165a7cd6b506b9ed5aa57fdd1d9a1e26

Observation c1785c1d-8c56-4343-a04d-f990d3fcfe49 · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

MOSS-Audio Technical Report MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.042011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:3a5c2c9ad6fa9677a9ca1b8bad3d5246320bc7e340d83ebeb389ca1aecb5e652

Observation 2fa4aeee-6e4e-44ed-b021-194bd073135b · outbound

This paper cites Joint Audio and Speech Understanding.

MOSS-Audio Technical Report Joint Audio and Speech Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.089571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:59f3248bb4d41faef118820f8e86eacc8031dd242e4a2dba8ff8797ab10fd032

Observation 8cf30ed0-c42a-477a-9c38-10e81a1c5d40 · outbound

This paper cites DeepStack: Deeply stacking visual tokens is surprisingly simple and effective for large multimodal models.

MOSS-Audio Technical Report DeepStack: Deeply stacking visual tokens is surprisingly simple and effective for large multimodal models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:8a1dc8e92b51b4d684b01eb37b29acfd6011f6086b201c6da6f6f9708fcc2887

Observation d7bcb2bb-69a6-48ba-85d5-6edbed107792 · outbound

This paper cites MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence.

MOSS-Audio Technical Report MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.050583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:8fed4af0f5f74c67dc9d8afd91486bf8fb9b5ea37102adb11081317eccfc28e4

Observation c9a4d7d6-2427-414e-a2be-34b3cf0d5bda · outbound

This paper cites Transformer Transducer: A Streamable Speech Recognition Model.

MOSS-Audio Technical Report Transformer Transducer: A Streamable Speech Recognition Model

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T13:12:13.684226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:23ffbf8b6004965250153fa0fec1163b51c664be288c1bccfda4c0aee146c7c4

Observation 29492b91-fa32-486a-bd0b-e733ab0bd189 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

MOSS-Audio Technical Report Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:338fb0497a594996b756f77cb95ac1ac96fedd8c74373cc4da48607be41f8d2a

Observation 82c14792-db0e-40de-9be4-3c637d99cb18 · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

MOSS-Audio Technical Report SUPERB: Speech processing Universal PERformance Benchmark

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:25.070239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:a20d45b3cc54bbfa0310c5c5560257dffc418c54b7b2692e81802e1e88be6a71

Observation 9155490e-5eca-45c3-b645-19c8d3a8447b · outbound

This paper cites Qwen3-VL Technical Report.

MOSS-Audio Technical Report Qwen3-VL Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.074549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:c3064ff3c38381e7efcbe6c222469a7fb6c40f32b0dde97612ed8bda622a537a

Observation c031f99c-4eee-4423-9a11-79cc20a86d16 · outbound

This paper cites MOSS Transcribe Diarize Technical Report.

MOSS-Audio Technical Report MOSS Transcribe Diarize Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:49.762568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:21a4308446578348d7a70d94b282d2784a57912e05c016abc74fae93b99264f6

Observation 85af66fd-6d64-46cf-82ed-a73c719d329e · outbound

This paper cites BEATs: Audio pre-training with acoustic tokenizers.

MOSS-Audio Technical Report BEATs: Audio pre-training with acoustic tokenizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:193201794494ab4d7d37621403964a208f4037c2cd59cb34c7fdbb76e7918549

Observation a674584b-8f8f-40dc-9bf0-fbe39c94a07e · outbound

This paper cites Effective Pre-Training of Audio Transformers for Sound Event Detection.

MOSS-Audio Technical Report Effective Pre-Training of Audio Transformers for Sound Event Detection

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.084991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:8de4d8352718941e556bc6d88bcd03c43571d0050454764a3901c482cd695dfc

Observation 991c5787-464c-4103-a247-2600900de08c · outbound

This paper cites Qwen3-Omni Technical Report.

MOSS-Audio Technical Report Qwen3-Omni Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.073662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:d61718fa067a6c3b1f9343ce99698900caa0ce81a4738168aecb5761cb183b64

Observation 435d8d96-5cc5-4fa0-bbd8-6286f3af106c · outbound

This paper cites Fun-ASR technical report.

MOSS-Audio Technical Report Fun-ASR technical report

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.020627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:b6f368eae01b6614270a9adf35b94913a57aecdf4ae818890f78e576dd3eb1e0

Observation 3540138f-e7be-41da-b26b-03a4953c4ced · outbound

This paper cites Qwen3-ASR Technical Report.

MOSS-Audio Technical Report Qwen3-ASR Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.081788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:6f065e20d64cf3683d999c63068c79f47fbb6012287da6cfe9d9d11ce02c3c0f

Observation 58e70b04-d5df-4ef7-be84-d9fd491a4d04 · outbound

This paper cites Bag of tricks for efficient text classification.

MOSS-Audio Technical Report Bag of tricks for efficient text classification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:bc0c1d567cc327fe7354cbfba12745ce5cee0068b70e4004b20575a7ff674e7b

Observation 180dc482-029c-4068-b5b0-9eb89575ada1 · outbound

This paper cites FastText.zip: Compressing text classification models.

MOSS-Audio Technical Report FastText.zip: Compressing text classification models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.024453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:96f025d81ac6bd215d89d326ccd37819e24acacd6821708a130d7ac6a20f827a

Observation d5aab964-c959-454b-9230-7efd00127a8f · outbound

This paper cites Scaling speech technology to 1,000+ languages.

MOSS-Audio Technical Report Scaling speech technology to 1,000+ languages

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:4b2d46fe0276452a958499d44e3e370b9940e8605a2194a6ea48c9dfed469525

Observation 89d43869-c876-4363-b6ea-e8f5e9620074 · outbound

This paper cites Scaling speech technology to 1,000+ languages.

MOSS-Audio Technical Report Scaling speech technology to 1,000+ languages

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:52372878354e8e023b8ab46288f5b704648ef2c1e3ca23a3f9fc29c7bd4dd974

Observation a03b87e3-1d1c-434f-a5c0-672390f50972 · outbound

This paper cites an unresolved cited work.

MOSS-Audio Technical Report Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:dbbd50d89b27cf208a3619abcfeb5d207945fbb183e8bf8cc3bf4fc612db546b

Observation 9cfec829-f9cd-47b5-b323-07c90c8155e2 · outbound

This paper cites Leveraging self- supervised learning for speaker diarization.

MOSS-Audio Technical Report Leveraging self- supervised learning for speaker diarization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:a44e3a77e9d9d60a504500d003cebd2522312189de152d7692142d42caf09c24

Observation 9c8970c1-e9b3-4385-96df-fe2f77c33b87 · outbound

This paper cites Xiangyang Chen, Shuzhao Li, Xiuwen Zhu, Yongfan Chen, Fan Yang, Cheng Fang, Lin Qu, Xiaoxiao Xu, Hu Wei, and Minggang Wu.

MOSS-Audio Technical Report Xiangyang Chen, Shuzhao Li, Xiuwen Zhu, Yongfan Chen, Fan Yang, Cheng Fang, Lin Qu, Xiaoxiao Xu, Hu Wei, and Minggang Wu

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:25.078694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:b93f3005b3100131eb671d1c5abf85eab74924304685920857a095abcc9e047b

Observation cabe4df8-1509-4859-98be-f8d12619485b · outbound

This paper cites Listening be- tween the frames: Bridging temporal gaps in large audio-language models.

MOSS-Audio Technical Report Listening be- tween the frames: Bridging temporal gaps in large audio-language models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.106858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:4500ba97e76507a62696e8c4a9072389856b6ea7dcc355a1afdf3724d69e34c7

Observation 3f1b8ab3-fdad-4dd4-b7e7-880150545f25 · outbound

This paper cites Bryan, Zeyu Jin, and Justin Salamon.

MOSS-Audio Technical Report Bryan, Zeyu Jin, and Justin Salamon

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.103097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:4170f41dcbf7a4af9a8371fd158e14130232a903ec4124ced7da47a448734883

Observation 544e3d97-bbb5-4a42-b50c-e3fec440bbf8 · outbound

This paper cites Music Flamingo: Scaling music understanding in audio language models.

MOSS-Audio Technical Report Music Flamingo: Scaling music understanding in audio language models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.014328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:081941bd07a2de54bce1d9ddec2331bb14505f13e623a25e7808bd7cec45899d

Observation 570a94d8-0279-42fb-b1c4-ed2ab0745ded · outbound

This paper cites Sakshi, Jaehyeon Kim, Wei Ping, Rafael Valle, Dinesh Manocha, and Bryan Catanzaro.

MOSS-Audio Technical Report Sakshi, Jaehyeon Kim, Wei Ping, Rafael Valle, Dinesh Manocha, and Bryan Catanzaro

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:306611040c62cd07834dc4a5e1d4fb90db696f5064677832ee2e76637614ff46

Observation f0d99664-cffd-4375-9bab-48efff57341d · outbound

This paper cites Approximate note transcription for the improved identification of difficult chords.

MOSS-Audio Technical Report Approximate note transcription for the improved identification of difficult chords

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:6a9ae3ff0b14b50b75460da3d8803b86c152f57d06828f1c787a3aef172349e2

Observation 13dc7cf9-557f-4237-953d-65435b0cbd77 · outbound

This paper cites Beatnet: Crnn and particle filtering for online joint beat downbeat and meter tracking.

MOSS-Audio Technical Report Beatnet: Crnn and particle filtering for online joint beat downbeat and meter tracking

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:7d69e1fb90e98229d722176ac1929a878c69bfa82dc8fedda0b1a9ca45a5310a

Observation 46cd78cd-555b-4402-b3fe-5dfd99eb7535 · outbound

This paper cites madmom: a new Python Audio and Music Signal Processing Library.

MOSS-Audio Technical Report madmom: a new Python Audio and Music Signal Processing Library

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.990354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:7f55c00ce7bab2601629cddb929d60e6e66ab1aff395e3f96791cc0ff5b443fb

Observation c5c39b61-9ab2-42cf-bf52-15183feb8203 · outbound

This paper cites Essentia: an open-source library for sound and music analysis.

MOSS-Audio Technical Report Essentia: an open-source library for sound and music analysis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:ec36c744848fe5f1c95c0dcdb7a7e398787c4caaeb1f9aa2388acdb7149b34ff

Observation 62dfaf62-8a9e-4991-8e20-170792b8da2b · outbound

This paper cites Recent developments in openSMILE, the Munich open -source multimedia feature extraction toolkit.

MOSS-Audio Technical Report Recent developments in openSMILE, the Munich open -source multimedia feature extraction toolkit

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T13:12:13.688720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:490b6879907e8babdd1a842163ff5a6314f6bf57ac3876d8bfbf0ab1e37a2ad7

Observation 96a6de44-74cc-458e-b677-f25f64246bed · outbound

This paper cites Codified audio language modeling learns useful representa- tions for music information retrieval.

MOSS-Audio Technical Report Codified audio language modeling learns useful representa- tions for music information retrieval

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:7a2efbf1328225cd33767f9eb23775bae535d609fd8d2a2190a5185768c8a7f1

Observation 892d7e75-acc7-493f-b0f1-01096d5d30e0 · outbound

This paper cites SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision.

MOSS-Audio Technical Report SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.984647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:87075eb5bb63e55cb51dfbd73673d4ce8b517e4f7026c6adc73381f8ac8932c4

Observation 731d1eee-03d8-462a-b9b5-bc1eae320375 · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing.

MOSS-Audio Technical Report SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:25.110474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:92c5ff05e9c6ee37d79fe6b224c4d57d493b214c15ea660746fdcb68a6e5fa52

Observation a3a7cb2a-6403-4ac3-afa6-4eb6e0ac3465 · outbound

This paper cites Unified Speech-Text Pre-training for Speech Translation and Recognition.

MOSS-Audio Technical Report Unified Speech-Text Pre-training for Speech Translation and Recognition

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.010703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:9426f92c856a3cfe81928c54fe10f1c78be30c31cfba306f37802d3c61c73919

Observation 3cc14cf4-5788-4e9a-b665-2d17d061b13e · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

MOSS-Audio Technical Report SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.093347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:f97a6dc265d1553918f3fbc1605cdd784bdb10ff8a425615b8424dc948de92a2

Observation 30a14c5d-4d17-4ed7-8736-72fb790a3865 · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model.

MOSS-Audio Technical Report Spirit-lm: Interleaved spoken and written language model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:453aadd82368f013d8bf24d3654e1de3e562a0beb9bd08b6de9d0c659fba742a

Observation 96592c46-965e-4138-bf16-40c353f910c1 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

MOSS-Audio Technical Report Moshi: a speech-text foundation model for real-time dialogue

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.997694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:e3fc690de7dcffe0416d3f2eeb0bd5bb8f61a0ebd15244dddd25a1477c747d31

Observation 314060b0-211a-4bc5-82a2-6dc28975b8ea · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

MOSS-Audio Technical Report Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.982340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:42b768af27451572364c6a58a5fdc939febdd6e8cd04b62580b6ee33124b0289

Observation 363b0d7f-2e34-4f19-b7e8-1c3d336660ec · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

MOSS-Audio Technical Report GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.986552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:e48273e3b96fa80cf276c3022d667a9f6265c037179f71c2df33bc16786fadf7

Observation 181163d2-d0d8-453f-8e9a-015a68694417 · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

MOSS-Audio Technical Report Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.114734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:eda85aeba2ad68dc94c984a083541692ef1c1574353e74d5b1c398072d5de721

Observation d4b3731c-80aa-49c6-84d1-d6fe04227c59 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

MOSS-Audio Technical Report Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.977576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:4952cce63c21ff22cd33478d35d690bc617bcde2171f89186b414ea8cddab8df

Observation 971d39a6-a5f5-47b7-909f-c7eb6657c1bf · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

MOSS-Audio Technical Report Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:c7b340c9b52255eaf1a00eddab2f86b0d731e9a9294a2f07fcad65afd5035269

Observation 3992ce51-9e1f-41ea-9f6e-6ae2ebcbacb9 · outbound

This paper cites CLAP: Learning audio concepts from natural language supervision.

MOSS-Audio Technical Report CLAP: Learning audio concepts from natural language supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:fb811304b76485e4378f3ab989c4eb86a87433a310d531365c247b4f76c74782

Observation 9cddd868-3b8f-4fc0-b0e5-10c8b458dd5a · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

MOSS-Audio Technical Report Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:b32110f664a3ca3e8f7d98e6e10cce8a0ada23059c25254d5b8080feae8877b7

Observation 7fb0502f-60fc-4dec-b079-65cbe03b1652 · outbound

This paper cites High Fidelity Neural Audio Compression.

MOSS-Audio Technical Report High Fidelity Neural Audio Compression

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.028612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:538beb18fefa6606fb61d23ff143d907168571e4e5626f4bd7fa2d95f13f9ce4

Observation 9483aace-6f76-4ce4-8758-b80fa5199566 · outbound

This paper cites High-fidelity audio compression with improved rvqgan.

MOSS-Audio Technical Report High-fidelity audio compression with improved rvqgan

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:b3785d625aac23fcbf24cc28b91125ace6754752414180095e819daf9d3ad955

Observation c02b7ec1-3b85-4b62-a36a-ec76c0345d7d · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

MOSS-Audio Technical Report SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.032216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:ce6417efb9afce7a40d824e9929b31c89d17d373d42d332ab097a7899cdb4888

Observation 1065d7de-b651-49af-9d75-51695c2fc376 · outbound

This paper cites Codec does matter: Exploring the semantic shortcoming of codec for audio language model.

MOSS-Audio Technical Report Codec does matter: Exploring the semantic shortcoming of codec for audio language model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T13:05:29.813707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:5cb450ad5a2ccf4b804ce8e141678ddb06e23c7847f3a93fc6f42f25f13f7e39

Observation 91131f25-74cd-4d44-9320-fcf040698176 · outbound

This paper cites SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding.

MOSS-Audio Technical Report SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.001719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:4d8fc281cf11109669b478ae03ddeda0307d73b545be79379f81b4da3ec6073b

Observation bec92328-9791-408e-9a7c-b75b49aadede · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

MOSS-Audio Technical Report TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.957516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:32fe1a9e105a94fd9a94ed9ddf237228704061e94cd94f48d011c9f651c7ce9c

Observation 0e53b4a1-9a58-4b03-aca8-204bec188cd6 · outbound

This paper cites The interspeech 2026 audio encoder capability challenge for large audio lan- guage models, 2026.

MOSS-Audio Technical Report The interspeech 2026 audio encoder capability challenge for large audio lan- guage models, 2026

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.099974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:b24a9e400890be4fd3b7fa6d47dab8a87694911ecc66aafa15b9dfd1eba016e2

Observation 85c0811a-7d11-4253-898c-821335d34d10 · outbound

This paper cites MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks.

MOSS-Audio Technical Report MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

Reference 66

Resolution
malformed identifier
local_arxiv, observed 2026-07-02T00:56:24.960610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:3904a5158291b361b2156169b78d4adcb18756f6bdfbf8b5702eebb795470be6

Pith citing papers

Observation 72f844c2-04dc-4cec-9949-df97b3a813a5 · inbound

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation cites this paper.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MOSS-Audio Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.831036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:9c5e3623bf774dc6a0618d4acc2d1f4e01d23dbcec87a4e5fc69cc48ba4a1250

Observation 765abb12-078e-4eb1-b78c-613e5d6aab34 · inbound

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing cites this paper.

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing MOSS-Audio Technical Report

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-07T14:53:55.823891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T14:52:35.127533Z digest=sha256:7a34e2e74fd107c775e0cf2b6b322d56eae951eca6603ce223aaffc8eff7aadf

Observation 259ddbf8-222e-4dfc-9a60-ca007a266265 · inbound

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing cites this paper.

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing MOSS-Audio Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T08:34:07.442930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:34:07.442930Z digest=sha256:bdcfa820bde1178f50c3f2d63358d973a9268183ba0b3011284a0ec03f3e4d65

Observation 1c3c4844-4d0a-458c-9cdd-2ee5f9de746b · inbound

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations cites this paper.

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations MOSS-Audio Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T19:08:12.244455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:08:12.244455Z digest=sha256:2efc14d68a2dbc40b397e31dc924215bc283548d6b26ce432385a33caee34f22

Observation 40ce1bda-0752-4599-90f8-53709420fa9e · inbound

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering cites this paper.

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering MOSS-Audio Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:37:32.397890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:37:32.397890Z digest=sha256:8cc22840a9767a8898291121b3642dd3d305d2a2cf47b2218f77b9fddfd6dc7e

Observation 4ab4d448-f56a-435f-9ae3-44456c282ad9 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens MOSS-Audio Technical Report

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:56.402445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:56.402445Z digest=sha256:231127aa10bbe000be98b957faceedc6b150629c8cd69861a11ae2b70a091913