Pith. sign in

Paper Citation Record · LEDGER

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 3 inbound Pith citation observations for arXiv:2506.00843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00843 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:43.372701Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:43.249937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:46:15.415472Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db0b3957-29f6-4b6f-b8cd-57d4aaa09ccb · outbound

This paper cites an unresolved cited work.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:43.870474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.240637Z digest=sha256:dd0894d7646db3932c53327688f660db65c7ce2da2bae81ec73d5fb157427bf2

Observation b4929345-79bb-4128-8beb-68d305901ca9 · outbound

This paper cites an unresolved cited work.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:43.860707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.244948Z digest=sha256:039f397fd1abc570c096f5f9aeeb3893a9d048617c3562a721b52f2b9e38303d

Observation 8bd3210c-2a8c-4609-9fee-b50842449518 · outbound

This paper cites HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.249937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.249937Z digest=sha256:824b9fe906300635baa496353ec8cffb757994ae77db8741085df5dcc1d9ebb4

Observation eeba0d67-d4b5-4d19-b4af-2ae36ba30a39 · outbound

This paper cites For acoustic training, we adopt the DAC framework [16], extracting random 5-second segments (vs.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement For acoustic training, we adopt the DAC framework [16], extracting random 5-second segments (vs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.851385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.253983Z digest=sha256:c36673308cc19afb2f08db464c037192b8296e67eeb370fe7ee935a7ed8a4b31

Observation 82947716-31aa-43ff-bbdf-3dc0d1a48e89 · outbound

This paper cites an unresolved cited work.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:43.840244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.258313Z digest=sha256:763d2f351552a2b83a1ce0c1c608352a0e0edc4012871e3521f4cff1e1589ce2

Observation 1e0b5dee-e1bc-4f0a-8c5a-9f05b98671d7 · outbound

This paper cites Our approach effectively preserves semantic performance for ASR while achieving reconstruction quality comparable to state-of-the-art neural audio codecs like DAC.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Our approach effectively preserves semantic performance for ASR while achieving reconstruction quality comparable to state-of-the-art neural audio codecs like DAC

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.828795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.262122Z digest=sha256:ae982f09eabd84934645d81f9b1f1866f411cbcfd79f3aa3bfd1349d3b6a0dae

Observation 0da11723-e0cd-4d2b-9188-999a55a0e457 · outbound

This paper cites Comparing Discrete and Continuous Space LLMs for Speech Recognition.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Comparing Discrete and Continuous Space LLMs for Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.265309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.265309Z digest=sha256:2874d89c4a072b9f2ad5b81fe0be098233b9876131fe993c9945df6fb83670a6

Observation aa2e66cf-500f-461a-a0b3-adc34cf2a18f · outbound

This paper cites AudioLM: A language modeling approach to audio generation,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement AudioLM: A language modeling approach to audio generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.816128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.268694Z digest=sha256:ab3c338110be0a3d6d89daa7749e802548794857db2632424829533deefaf002

Observation 2693840c-85e6-414a-85aa-b7687e230ab6 · outbound

This paper cites A survey on speech large language models,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement A survey on speech large language models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.272054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.272054Z digest=sha256:a3139a662ba3db23fe7934148e2f2616193ae3bf63d7924fc6b809a99add001e

Observation 99b8b4a1-1f61-4892-aff3-7334e25a17f6 · outbound

This paper cites SpeechGPT: Empowering large language models with in- trinsic cross-modal conversational abilities,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement SpeechGPT: Empowering large language models with in- trinsic cross-modal conversational abilities,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.800601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.275346Z digest=sha256:034fd365fea8e6f0d31fd43045c17d2dddc2ed2a8a096c092227f73bdccfa28a

Observation 8267a6c4-ff08-4903-b453-4e3fed29d9db · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.278401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.278401Z digest=sha256:5ad6f57e19d57c3942e2ead0459e85ba112be1c3c520f86d585a13f6659525b6

Observation 7f951cb4-542c-4b3a-b069-28b70dc236dd · outbound

This paper cites Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.788909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.281631Z digest=sha256:85f5bb5058926fe69cf9a87b0765e6f4f3d442a39ca8f412f0f64331e01a103c

Observation a28f4ebe-6fac-4bd1-9202-ff7f9a27a104 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.779218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.285176Z digest=sha256:cf009bf2f9fff67c7153b9dba8e8c39950c358bdb23402e4d9782fb37204d1cb

Observation 62870af1-70b1-4e6b-add9-e5cae4802663 · outbound

This paper cites WavLM: Large-scale self- supervised pre-training for full stack speech processing,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement WavLM: Large-scale self- supervised pre-training for full stack speech processing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.770201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.288493Z digest=sha256:932f0adaca55bff484fcae867b01c256d4cdb7500f13ed10e9c0b787ebe8cff1

Observation 39eac272-4f3c-464f-ab0c-9e0b666e396d · outbound

This paper cites W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre- training,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre- training,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.761050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.293070Z digest=sha256:2644eadf950be3091c5e4cdcc9aea50a24dfda3eb2947bcba60456cff6aa16a3

Observation b7bc03b7-2eb0-481f-9ebb-67c954e0e0d9 · outbound

This paper cites Ex- ploration of efficient end-to-end ASR using discretized input from self-supervised learning,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Ex- ploration of efficient end-to-end ASR using discretized input from self-supervised learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.750837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.297026Z digest=sha256:c3af70790a48e15c890fcd323c836ae3f96b921c9d3431c81c4165558a311962

Observation 63762711-6ffb-4f12-9c90-beb58a970dd6 · outbound

This paper cites Speech resynthesis from discrete disentangled self-supervised representations,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Speech resynthesis from discrete disentangled self-supervised representations,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.740820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.300050Z digest=sha256:2b2f9bb70022e6f311b25e59ea9713de0c13d1f7342bb45f43fa7879530a4479

Observation 1debd9fd-98c2-4a19-af30-efeb8f8a53c2 · outbound

This paper cites Ex- presso: A benchmark and analysis of discrete expressive speech resynthesis,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Ex- presso: A benchmark and analysis of discrete expressive speech resynthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.731052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.303897Z digest=sha256:88d5871ba279aff8e2a851920f217c4248128d5933f2cd9da4e3a1291f3c18ce

Observation ad690667-0ceb-4340-948d-a5881e8ed5dc · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement SoundStream: An end-to-end neural audio codec,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.720715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.307670Z digest=sha256:a822ab6053d0d3024db2d666c8e2d7361590557c5cbdaf20dd115b6a72ec6ed9

Observation 14e5435b-bbc6-4a45-84ca-5fc60c5ad807 · outbound

This paper cites High Fidelity Neural Audio Compression.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement High Fidelity Neural Audio Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.311743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.311743Z digest=sha256:bcb7c885d5ee2454b6ac74de600d3d184d8122bb73c7179c4fa0bf15f836077e

Observation 1bec4980-6a30-4924-a5a5-417b5b962cdc · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Moshi: a speech-text foundation model for real-time dialogue

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.315618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.315618Z digest=sha256:967c022987ffcc860304860dcf9548a2250b0325feb9112f2ac64db1d7d29e7e

Observation 69d89b07-bf94-43c7-b1d9-5e91394365e4 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement High-fidelity audio compression with improved rvqgan,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.708488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.319166Z digest=sha256:00405d75a1c69dfab11de8793b96fed12a9dc29b64bbc87c8203550e8a01315e

Observation 5b674017-f082-4303-b5d2-94a79dfa57b3 · outbound

This paper cites Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.322223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.322223Z digest=sha256:432cbb5121d03863ca944f6fa1343340797d2042a4e7d827f96ae08a2d6e653f

Observation 2c5d41c7-f072-4f97-8fc0-eb203deab3ff · outbound

This paper cites QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.325585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.325585Z digest=sha256:3b48b3848928042ecddcdcf6d43c07ba3f6098b7411d576e0815f85842df0830

Observation ac8b6874-1279-488e-ab19-fcd25f5d331f · outbound

This paper cites MMM: Multi-layer multi-residual multi-stream discrete speech repre- sentation from self-supervised learning model,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement MMM: Multi-layer multi-residual multi-stream discrete speech repre- sentation from self-supervised learning model,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.697823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.329456Z digest=sha256:1cf20adea039c4f25b4673a8b237b6381aaa00bae9df35585e3fe4d4bf9b0476

Observation 9550b58b-0ce6-43f0-9933-47545d4125e1 · outbound

This paper cites Towards universal speech discrete tokens: A case study for ASR and TTS,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Towards universal speech discrete tokens: A case study for ASR and TTS,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.687447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.332570Z digest=sha256:7414a07041a96f9bb718bd3b469e63bf83ddcc2cdda9c1c12edd48c950d34506

Observation 5229250d-6dd1-4705-8258-566427df1879 · outbound

This paper cites ReVISE: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement ReVISE: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.675974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.335931Z digest=sha256:b4d835b18f705e11cc9e23c59d8858e623e8593e1a14fd193225c01e2707eea2

Observation c93f827d-aab7-45c8-ab08-d67d4ec6893a · outbound

This paper cites Self-supervised disentan- gled representation learning for robust target speech extraction,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Self-supervised disentan- gled representation learning for robust target speech extraction,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.664679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.339284Z digest=sha256:f754f8be2aded8b2aaca1318e6e3108de986b9d29e559c9fecc7252a13a6d2ed

Observation 5457c13d-fa21-4af3-aece-434aff075c76 · outbound

This paper cites A ConvNet for the 2020s,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement A ConvNet for the 2020s,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.653509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.342683Z digest=sha256:11f629f3f05c611f1c26368eb92966a36b0d3ede3004a44735f0ee75a45d4565

Observation 1c0dc1e9-17c3-4a52-841d-b2ddf644724c · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recogni- tion,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Self-supervised learning with random-projection quantizer for speech recogni- tion,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.643567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.345879Z digest=sha256:0e654a1cdfe1643170d49e7d97f233d701881e665c9207bc31790c3b07c4ac96

Observation 7cbdc036-e0b1-4139-b856-a6d2a7927a1a · outbound

This paper cites Mel- GAN: Generative adversarial networks for conditional wave- form synthesis,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Mel- GAN: Generative adversarial networks for conditional wave- form synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.633237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.349431Z digest=sha256:1694c66b196b601b5c202dfb96f8b1256aa285621c56bdb62e6bf3b272cbcb52

Observation 46499834-a384-4020-954c-719ab49dde7d · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.353811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.353811Z digest=sha256:e1ddfe6ad8ddc0c4547a7f70f61571d0fc591e34d8d41946abf3508ee1b20b31

Observation 8f28f306-a4f7-4dee-a2c4-9d2e8bdd0af8 · outbound

This paper cites Lib- riSpeech: An ASR corpus based on public domain audio books,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Lib- riSpeech: An ASR corpus based on public domain audio books,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.615831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.357062Z digest=sha256:817a83b514ec76f80ade8940996ce8a2aef7606a0b7dd694298d0bb6b2c47499

Observation 13b65055-9954-4a90-baa7-7cd5312b1987 · outbound

This paper cites Lhotse: A speech data representation library for the modern deep learning ecosystem,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Lhotse: A speech data representation library for the modern deep learning ecosystem,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.605540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.360747Z digest=sha256:3fe19eae6aa9ab9167f85d06f2be1ce7094e0e7507855a4bf410d373aba06844

Observation e29dcf14-9587-42e5-ac29-b1656da52e00 · outbound

This paper cites Open Implementation and Study of BEST-RQ for Speech Processing.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Open Implementation and Study of BEST-RQ for Speech Processing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.364043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.364043Z digest=sha256:375b84e64c023f1278b429cab95e9e63acf02c747233b989cec79736590f2ad9

Observation a8a9fc1d-9240-4912-9dd7-2837dee109f3 · outbound

This paper cites Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.595072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.368593Z digest=sha256:c576dfcaa09e1289bc018fa81971d9cae9ecad0acb269f3009c2d88672d3fe36

Observation d64d0e91-ea06-4e9c-87b0-c57f1cf42165 · outbound

This paper cites ECAPA- TDNN: Emphasized channel attention, propagation and aggre- gation in TDNN based speaker verification,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement ECAPA- TDNN: Emphasized channel attention, propagation and aggre- gation in TDNN based speaker verification,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.584006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:43.372701Z digest=sha256:e9bcfe8c0321e6ebb6b804c4d47224eff98267ffe154f22d27af6976ba0536c2

Pith citing papers

Observation 8bd3210c-2a8c-4609-9fee-b50842449518 · inbound

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement cites this paper.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.249937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.249937Z digest=sha256:824b9fe906300635baa496353ec8cffb757994ae77db8741085df5dcc1d9ebb4

Observation da8f0ff8-c3eb-4a88-8fbf-6acdd6e5f6ab · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.417161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:01d90be90c6235c4ba9054ccfcc8830b543fa4e9063b964fd2c524c615eec335

Observation 98a9678c-94cd-430a-88af-17b774c73b52 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.653335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:9a4abe28793f22c425f2f0cf579b88522dff00fe3a42cdcfba087fe2b85790f2