Pith. sign in

Paper Citation Record · LEDGER

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

As of 6 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2603.05094.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.05094 v3

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T15:40:07.421894Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T15:40:07.421894Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-15T15:41:12.087219Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact12
  • verified fuzzy29
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2ad3e17-146d-47eb-8c94-d0335a3f926a · outbound

This paper cites TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:41:12.090471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:6f386abfa02b1e3e751e6b6083a9000c7f62a6f2f334bba37c9e3817aeee03f1

Observation c1d7ee02-2725-4d0a-b589-b43a669cf98f · outbound

This paper cites an unresolved cited work.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:41:12.535746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:8f787931f1c264c8559b507bf6696463ce8e387ff35559f93fb12a4f24d751df

Observation 58dc4f1c-1a71-4e0c-bdd4-4c66a4cf9c8a · outbound

This paper cites acoustic long-tail.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling acoustic long-tail

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.539993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:b9b7ec8f78c173623c9e501e9b9f80377d317879b87f109bbf8892afb43cd88e

Observation c3a18b3d-3762-497f-a18e-4ec2015e42aa · outbound

This paper cites an unresolved cited work.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:41:12.666275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:b9f58aba0d7bc4ac12033d3b5ca94cad407f58063e41842e5157c47c02f67b69

Observation eaabe514-74a4-4f48-ace0-10d9ab9592d6 · outbound

This paper cites To preserve speech-free soundmarks, clips where both ASRs yield empty outputs bypass the text check.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling To preserve speech-free soundmarks, clips where both ASRs yield empty outputs bypass the text check

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.656907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:07e5616c00591430b7974d931b67b9a2be45346700e214211b1c12601307c98a

Observation a5c7abf1-20f0-4e0f-b76f-ef565e77dd8e · outbound

This paper cites an unresolved cited work.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:41:12.661985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:e1ef0fa7f3a7e96f71ae8467b852e7eb9e4ce0f371358d3c5026df6ab6503b37

Observation 61b305de-86fc-4728-85e5-ce1b2c649505 · outbound

This paper cites Hello everyone.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Hello everyone

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.670278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:6b87fdd33691027175b788a15c87d15c8f639c01d937a9ee1031e4a808084113

Observation ba5f6592-67c2-4563-a3e3-aafa29c606b3 · outbound

This paper cites Experimental Setup Implementation Details:The proposed model, Tai-LALM, is developed as a localized adaptation of the DeSTA 2.5-Audio framework.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Experimental Setup Implementation Details:The proposed model, Tai-LALM, is developed as a localized adaptation of the DeSTA 2.5-Audio framework

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.544488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:33b6b0bdd00f291cb6763f523d15fbdd89bf1f666d9713349f49172eef1e5770

Observation 944ae7f0-7b37-4398-97ba-a9f7769ac5c6 · outbound

This paper cites Architectural scal- ing alone is insufficient for robust sound-to-meaning ground- ing without localized acoustic semantics.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Architectural scal- ing alone is insufficient for robust sound-to-meaning ground- ing without localized acoustic semantics

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.643132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:295a215da1fc63c508ffc7cf432a2762c0e1669e252b620008677b6576461999

Observation e02ad1cd-4a20-4dcf-8ac0-0ab1f5aaf742 · outbound

This paper cites Beyond the Taiwanese context, this pipeline offers a method for regional adaptation.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Beyond the Taiwanese context, this pipeline offers a method for regional adaptation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.647763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:7d823d89d9da91daeb4aa9829669e11f2e9914d8adb29dfaa133c7f875a81803

Observation 3ab9292e-e197-4519-b3b5-3d214061f0da · outbound

This paper cites The perfor- mance gains underscore the necessity of the VGC pipeline for robust training-time curation and Dual-ASR arbitration for sta- bilizing inference.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling The perfor- mance gains underscore the necessity of the VGC pipeline for robust training-time curation and Dual-ASR arbitration for sta- bilizing inference

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.674762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:58277e64a15033cc6c98e654975d8307ea50a5faf93c776f4cde6db26963dfc3

Observation b65cae82-50fa-4474-8a71-11e63c0c172f · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:07:57.965390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:f9bd6369e91ca44a551c27488b39448e8991de9344c11aa2bc9b26035edb5099

Observation 75e9f46a-c54b-45ad-b32d-60bdc1d41b2e · outbound

This paper cites Dynamic-SUPERB Phase-2: A collaboratively expanding benchmark for measuring the capa- bilities of spoken language models with 180 tasks.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Dynamic-SUPERB Phase-2: A collaboratively expanding benchmark for measuring the capa- bilities of spoken language models with 180 tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.633129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:f797c7df049751017d8dba50b71411a30045b969827830f9c41f69a2ad5e56bc

Observation 8dc8a962-a94b-49f4-a55e-d0b1629820e1 · outbound

This paper cites Listen, think, and understand.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Listen, think, and understand

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.552340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:9cbb46b36dca86596c2cb2325be27ced061724cbd150060a8b2c1eb21c807219

Observation e0a623f6-c828-4464-b9e0-9fc2fc9bb675 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling SALMONN: Towards generic hearing abilities for large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.565753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:e7dfd193f0495920917c36cff93c92bb286f6346bfd143095ba63575a673855e

Observation 4eaee3f0-f2e7-4f2a-89ec-0f84c995de76 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:41:12.014149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:677e2820136a36409dddd6586e2e2dda99312849bf3ce2531349f68d0f7de786

Observation cce735d0-dced-49b6-86da-d08e9df6816c · outbound

This paper cites CultureLLM: Incorporating cultural differences into large language models.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling CultureLLM: Incorporating cultural differences into large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.548439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:579c95edbe7ea70dbb3727f0c812b72f7ce5a4e39dbde0431b97af3818e18486

Observation 3adab25d-ddcf-42b0-b2ae-5dffbb57b0ea · outbound

This paper cites Universal paralinguistic speech representations using self-supervised con- formers.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Universal paralinguistic speech representations using self-supervised con- formers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.624284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:2657190fa25cf94a5925f7e7e2c770da47a8d4bc0cfac1ee9e75ba1326261ebc

Observation a00cccf8-d019-448b-8617-d4d74dfe1569 · outbound

This paper cites Building a Taiwanese Mandarin Spoken Language Model: A First Attempt.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Building a Taiwanese Mandarin Spoken Language Model: A First Attempt

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:41:12.036461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:8a4dd1666a4d9a257d87d0a621009d139c3a9673123800747111e99e49696424

Observation b4a85159-a69d-4ed1-a07f-9f381714c7cb · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.611522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:e1e3751d2e10f2df04b161a1e4fa3d9431cc6e37f82066caa45434b1f472d2fb

Observation 04407e70-5fd3-4939-afd3-cddb23e862ff · outbound

This paper cites Mitigating subgroup dis- parities in multi-label speech emotion recognition: A pseudo- labeling and unsupervised learning approach.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Mitigating subgroup dis- parities in multi-label speech emotion recognition: A pseudo- labeling and unsupervised learning approach

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.556825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:2a67be9f40c420f1744e57b886be83ef14a4ebcaf098a85f3d34e573445cb339

Observation 0110998a-97a3-4756-a826-22975c002534 · outbound

This paper cites AudioSet: An ontology and human-labeled dataset for audio events.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling AudioSet: An ontology and human-labeled dataset for audio events

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.615657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:4509bd7321e8f8b986c31f8450c2fa7a9287769e5b41546c72701afd8ebb3037

Observation cd723bb6-25f5-4c47-808b-76a4968ca288 · outbound

This paper cites Lib- riSpeech: An ASR corpus based on public domain audio books.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Lib- riSpeech: An ASR corpus based on public domain audio books

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.561496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:b608757c50f902f41e619d4be26cba70953b6a1ad1ff32e810ecb44fdd82c1ac

Observation 31f3d6d9-cfe5-4661-a66a-4cf0dea0c4cd · outbound

This paper cites WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.652162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:ed136714df565cb9657ff545604a76fcbf5365df9adc1009544eb6d44693a437

Observation 9bc4780a-01bc-4901-8c04-2fb1c7fb6497 · outbound

This paper cites AudioGen: Tex- tually guided audio generation.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling AudioGen: Tex- tually guided audio generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.607808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:fce14b08b2cadda4d4d904714f94a183796c14b7eb9932f62fa7414549060a82

Observation da662ac1-43f3-41c4-bf05-78d6facf5437 · outbound

This paper cites When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:41:12.029178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:463617392a790d4cb4452358bebb52e0682cadea0be49557b6cae0bdc3039096

Observation 90cd461a-39cf-41a4-a45f-a9b9926f70e0 · outbound

This paper cites WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:41:12.042921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:a96441bb8e3cecc59dcf79a3deede2a62d2b97d01f7c6a4cd1326163b49cb026

Observation b8819a61-e5ed-4733-8bf8-cacc9813cd48 · outbound

This paper cites Ke- Speech: An open source speech dataset of Mandarin and its eight subdialects.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Ke- Speech: An open source speech dataset of Mandarin and its eight subdialects

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.569860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:6b40b5eb99fd63436cc4f0342837828338f9a9e3c2c8f40cc8367c57fa350a8e

Observation 0863d271-bea5-4dde-8292-44d96213b520 · outbound

This paper cites LESS: Large language model enhanced semi-supervised learning for speech foundational models using in-the-wild data.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling LESS: Large language model enhanced semi-supervised learning for speech foundational models using in-the-wild data

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.597404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:70d8f6ba8346c56c3ada723018cff21b9fd06533af11950548f705c25d5e2a23

Observation b00ad799-b7d4-4046-a998-7c2f3cda719e · outbound

This paper cites Training language mod- els to follow instructions with human feedback.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Training language mod- els to follow instructions with human feedback

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.586867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:d90d704c6bd0bc5ad78d15a5714402ecbe377470d76a78c62daa4a31ae2bcae9

Observation efa01878-ef1b-4e10-8478-cd0d23bff92e · outbound

This paper cites Data-centric lessons to improve speech-language pretraining.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Data-centric lessons to improve speech-language pretraining

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:41:12.056262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:ba3a82338a0c4d550a9cc56516a7c567d9c54e17f928002b48ea737aec259471

Observation 300b3fd7-8048-4c67-ae5b-09cced34f132 · outbound

This paper cites Reducing ob- ject hallucination in large audio-language models via audio-aware decoding.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Reducing ob- ject hallucination in large audio-language models via audio-aware decoding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.592199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:89c475871638ef81c0ddbaea2dce058f48b62e704b104e562ee2b259a26e5549

Observation 075cd9f5-c2ee-428d-8f81-d231c6423f2a · outbound

This paper cites Beyond transcription: Mechanistic interpretability in ASR.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Beyond transcription: Mechanistic interpretability in ASR

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.602913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:8a631369870968ecbdbcfbb846353a13507509298e5f58dec5a9e7d4feaf4302

Observation 46975b0f-9fff-4e39-bb9c-6e1e06c911ac · outbound

This paper cites TAU: A benchmark for cultural sound understanding beyond semantics.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling TAU: A benchmark for cultural sound understanding beyond semantics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.619954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:34a03c6e5a1ca0af3035985e59e47289aa361c58753dad10e75ba495f0f63646

Observation 63112cbf-69af-4d20-b282-69d48e16c237 · outbound

This paper cites DeSTA2.5-Audio: Toward general- purpose large audio language model with self-generated cross- modal alignment.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling DeSTA2.5-Audio: Toward general- purpose large audio language model with self-generated cross- modal alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.628540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:bd82fd40a73a1eb4ca724c2106ff24928c4af8a4b2ef72e61679c42ea8bd16a2

Observation fb14dec9-9e30-49d3-ab98-5dc0eda6dfb4 · outbound

This paper cites Attention- passing models for robust and data-efficient end-to-end speech translation.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Attention- passing models for robust and data-efficient end-to-end speech translation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.638215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:a7020aab98efb0958013ba19222205f66e183e445ebcd921ba41d15e1aa19282

Observation df6cac8d-acff-4af8-b471-8312b88f3a36 · outbound

This paper cites Lost in transcription, found in distribution shift: Demystifying hallucination in speech foundation models.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Lost in transcription, found in distribution shift: Demystifying hallucination in speech foundation models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.582394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:91c406569d0920717dfad74191f45524739c8d35d566c768ea6fb89517ba27e0

Observation eafd01da-5065-411d-8b0c-f6a289d81f4a · outbound

This paper cites Teaching audio-aware large language models what does not hear: Mitigating hallucinations through synthesized negative samples.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Teaching audio-aware large language models what does not hear: Mitigating hallucinations through synthesized negative samples

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.577996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:58a0e1fb89908dedb310c01cef361df38e0914149d84ccbe6fb829f7e9da4af0

Observation 2086a53c-cd28-4c44-8134-1640c39e5fea · outbound

This paper cites Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:41:12.022755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:c5afe1dfc2e986692e2cf69f83c61ad9724fb8d567b7ee76bf6b2f4d6372b2f6

Observation bd2de4d1-4754-464e-be92-2f2d992f940f · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Robust speech recognition via large-scale weak su- pervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:41:12.573996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:88b798f1058f55b8ee5e86f082039e794773a470fe162cbfff40e0d1ebe709da

Observation 814d604a-bab6-4ee4-a3bd-e5d0fba97762 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:41:12.084111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:76c58f0c4695f091d74b4eefdc7fc5363faeffe6c8946a0d967e52a858e3a2e8

Observation c05e7224-bbf4-4475-ab49-064a0f8c7e9c · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:41:12.062679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:6bb13ea3492e45cc679f8eea64b2e6d71c81ddb08e5b5f4a9667cc6626e70051

Observation d1ab2d00-a676-43e9-bbfa-04079f2a6d73 · outbound

This paper cites Qwen2.5-Omni Technical Report.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Qwen2.5-Omni Technical Report

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:41:12.049671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:99646d1da618aee14732e21782516cdfeab689a434b5a6d326830abf4ff5738c

Observation 33308ca8-b00c-4034-be1d-b4d7af0aec01 · outbound

This paper cites Qwen2-Audio Technical Report.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Qwen2-Audio Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:41:12.070080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:0a553445a6303a0757ea45d6b3c7d0d119a9b0dcdedfba5f1afca989e3c89274

Pith citing papers

Observation c2ad3e17-146d-47eb-8c94-d0335a3f926a · inbound

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling cites this paper.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:41:12.090471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:6f386abfa02b1e3e751e6b6083a9000c7f62a6f2f334bba37c9e3817aeee03f1