Pith. sign in

Paper Citation Record · LEDGER

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

As of 17 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2507.20091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20091 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:19.340049Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T00:49:26.507281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:46:15.325325Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dee9aa13-00b1-4644-b147-bbf0eba5df4f · outbound

This paper cites write newline.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.095238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.095238Z digest=sha256:87e06f08704bec1e1cdad2898735b75e386fa1d3a99c84d379165563aa0ef134

Observation 7d4af0ee-20f8-4091-a709-a63d1d3f8468 · outbound

This paper cites Dm-codec: Distilling multimodal representations for speech tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Dm-codec: Distilling multimodal representations for speech tokenization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.101576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.101576Z digest=sha256:fcfd00aaad69f55d2fa75031c5f943c7a1a1fb6c909b2b5daa5a23a61c410720

Observation c92597f9-8e37-453a-b3f5-2f29d8379e52 · outbound

This paper cites dMel: Speech Tokenization made Simple.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models dMel: Speech Tokenization made Simple

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.106677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.106677Z digest=sha256:53c213e25e00cd7a9637257e727e8f90eab4da095ce9709aede0f642173179ce

Observation dddf75bb-7c71-40a6-9544-69c2e8c066c2 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Audiolm: a language modeling approach to audio generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.112745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.112745Z digest=sha256:c885663b899573d27aa69856523bc0953d85794f34c517fba841d1c91ca9b278

Observation 6d818c27-e966-48ff-85f2-dccce0497ae0 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SoundStorm: Efficient Parallel Audio Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.118430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.118430Z digest=sha256:8c9eba13b87d76456e6aa40a4e235e5770f8ffe3902bbc6a0c6343946313589b

Observation 0e085d97-ffc2-4cef-b51c-8f5fe7d5dae0 · outbound

This paper cites Giveness, contrasitiveness, definiteness, subjects, topics, and point of view.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Giveness, contrasitiveness, definiteness, subjects, topics, and point of view

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.212182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.123781Z digest=sha256:8c17e32d13888f3c132990c362000202bd8dec55adcdb690be79ce37b1244f23

Observation 3d1e0e8c-5558-49b9-85ea-5876e630df6e · outbound

This paper cites DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.129079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.129079Z digest=sha256:5726ca036f6a1e4eee74e7242ee0c3b40709647055fa9da25999089b46f0b687

Observation ccb609e8-ddf5-4482-843d-273d1b8e377a · outbound

This paper cites DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.134392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.134392Z digest=sha256:78d1b9f64f19a949bf8d968d4919fd5c9807eb6d116450122e3c5afaf5c29bde

Observation 24823737-1366-4758-af51-1ecf7adb7e91 · outbound

This paper cites EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.140208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.140208Z digest=sha256:06c3dedc8ef085adb88ed02537cfc575a3c26844cd7de542802daccfb1dff93b

Observation 713a6d3e-8e89-452f-9f2a-83960024e07b · outbound

This paper cites High Fidelity Neural Audio Compression.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models High Fidelity Neural Audio Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.145484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.145484Z digest=sha256:cd4311dd09bbebfa42978438022822435dc06e497855c83b3201ffb5a13af406

Observation 7fccdd45-b965-454c-8f5b-97c77fc1e9f2 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Moshi: a speech-text foundation model for real-time dialogue

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.150827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.150827Z digest=sha256:5860d4bb5d9cc2b3fed35a3520abfb0fc94aac0cfb17cd78d001d3f54d8fec77

Observation e274388a-c2c9-48b9-8754-7b5b1eedbe3c · outbound

This paper cites Elevenlabs voice generation platform.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Elevenlabs voice generation platform

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.195688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.156119Z digest=sha256:8484fdd55cee41db4111b9c234a014d6d819f927e8af83b34e192829c938ed81

Observation 5612812a-b60a-4f0a-9d37-dae60a462b24 · outbound

This paper cites Recent advances in discrete speech tokens: A review.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Recent advances in discrete speech tokens: A review

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.161618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.161618Z digest=sha256:7c2977057b072c256478a212a1eb161e4ea409f8927affc35ba9b088489d119d

Observation b0d96cdc-cac7-4247-9e64-4290d0ad8f7d · outbound

This paper cites Lora: Low-rank adaptation of large language models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Lora: Low-rank adaptation of large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.166411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.166411Z digest=sha256:63c331b6dc03586e8dce318398b0eb5c3ddec4d698514725ade43f3480e420f7

Observation 285c13b6-7cfc-4155-a2fc-72d0f0da6136 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.171625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.171625Z digest=sha256:e787d8e0f7a0bb6d33a641d0ae45b963b3d78b1898cb8b2ba7f8dd6ba3630c6d

Observation b2f588c9-eb8b-47b3-8819-c41769fd9cdc · outbound

This paper cites RepCodec: A Speech Representation Codec for Speech Tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models RepCodec: A Speech Representation Codec for Speech Tokenization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.176618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.176618Z digest=sha256:a76b8bb9a137915499437cfabc74aaaba4cc7b890945d2eac59e86e487a7de15

Observation 73502204-a433-4ace-8bee-40ce7176eac5 · outbound

This paper cites Crossing the uncanny valley of conversational voice, 2025.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Crossing the uncanny valley of conversational voice, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.170020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.181649Z digest=sha256:8fe29b031d74c4042bf49b58b4405aee8c73d11f8205b638cd3e8fb7837ca2ba

Observation 357c6328-21d4-44b9-a094-d50b74a4d917 · outbound

This paper cites An open source emotional speech corpus for human robot interaction applications.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models An open source emotional speech corpus for human robot interaction applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.186416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.186416Z digest=sha256:d8e4dbd2658d9f2523b9afa60c8a279e13c55e57ce7f362d584037db36deb451

Observation e25b62d3-fec4-4b4c-aa00-4d0ed08db09c · outbound

This paper cites Style Mixture of Experts for Expressive Text-To-Speech Synthesis.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Style Mixture of Experts for Expressive Text-To-Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.191217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.191217Z digest=sha256:1de25652af720848cc62a79b6ad41fee7437de32c9667c267acea643130fa0ec

Observation 4c2ec780-f37a-4266-a023-2e9471ff796d · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.196492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.196492Z digest=sha256:c5193f4c2d853c53365a452622b1bfd03a15252b835f64fd45a9d17359d7c74f

Observation b7da3462-7686-40ce-81f2-9586298d6314 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Libri-light: A benchmark for asr with limited or no supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.201637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.201637Z digest=sha256:c069ffa397104f016f44a87eb9940996024cfe19099a4a9d4947ccb1b4bb2218

Observation 8995600a-5152-4815-b6a7-4fd51b794054 · outbound

This paper cites Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.206278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.206278Z digest=sha256:40990baefe8ab64165be9ea44c8e98d8c8791b2bbee3c86e45de145c8ac22dff

Observation d243f73c-5256-410b-8999-06111fd74af7 · outbound

This paper cites On generative spoken language modeling from raw audio.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models On generative spoken language modeling from raw audio

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.211165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.211165Z digest=sha256:18e5235f91d46878760d4e03697a7fc85049f94bc5617ddfc0eb156665162b78

Observation b6e6aa01-e127-43e4-b570-f2ed7cc28c2f · outbound

This paper cites Whisma: A speech-llm to perform zero-shot spoken language understanding.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Whisma: A speech-llm to perform zero-shot spoken language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.125620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.215741Z digest=sha256:fd73f7563e15822b95ae79910d703ad754c8be562ae8ea3f064aa45bb5fe00f8

Observation f5543cb5-b5ab-4813-95be-6a0a387dfe0b · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.109798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.220379Z digest=sha256:51f707f601361b84a56322c2793c225b3b136f38f8723ec21d0e839ebc0d9dfa

Observation dae726de-a84f-464e-a818-55039915a068 · outbound

This paper cites Generative spoken dialogue language modeling.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Generative spoken dialogue language modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.094079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.225998Z digest=sha256:77cc9a1e062e79ba18f9c45f91aabd322fcd83d8bd97d4cbffefb90e2a866752

Observation f05a6762-2cf0-4e7b-a310-5d3613d603ee · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Spirit-lm: Interleaved spoken and written language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.078237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.230893Z digest=sha256:24ae00bcca12c471759aa9795a2ec592700ffcfe223a50d4ce1bf69fdaf4707a

Observation 1da4e900-f7a8-4349-9c57-8b093de1d1f8 · outbound

This paper cites Long-Form Speech Generation with Spoken Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Long-Form Speech Generation with Spoken Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.236403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.236403Z digest=sha256:cb2f9be534ff94b758bdd855d6820641c914ec4daeec13b196721f623e6a7ef3

Observation d9db9bdc-d2cf-4d6a-bbd2-6002769aedf8 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Robust speech recognition via large-scale weak supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.241198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.241198Z digest=sha256:573239c5a60994b8705533b85803586f4e99c77d808584f40767def3f8c28342

Observation 912c46e5-47a8-4262-a262-cbcd8c6d09ab · outbound

This paper cites Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.052853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.246565Z digest=sha256:7e31651d89c7773f499a681ec96fa4b00ee9f21fffec8e54103de78ff5e5bcbc

Observation 8c198c27-1d04-46e8-90d3-e93b341966b3 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.251401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.251401Z digest=sha256:bb824deb730b4077e321fd54a76c870202f63bcdd1d31c457334638276624d12

Observation 57f1eb8c-049b-4b88-b50c-66b3bf03ffe1 · outbound

This paper cites Shechtman, S.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Shechtman, S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.036673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.256580Z digest=sha256:586b9846961adc7e194a6b3c2759fb386a234f07c2f54500119296ad2fd18a33

Observation 6fd81e2d-7668-4d24-bc95-5e4ec3d86d5b · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.261079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.261079Z digest=sha256:75be31050b29993fd2f0effd2342c0f6faefe3ca0a98a8ac607a1a03e381d650

Observation 5d103bf7-035d-42bb-9beb-54c4f55b5afa · outbound

This paper cites An analysis of the use of qualifications on the amazon mechanical turk online labor market.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models An analysis of the use of qualifications on the amazon mechanical turk online labor market

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.020204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.266649Z digest=sha256:e747491ad77eb2ff04fc50838b3523253250ec61112dde4fa7a983845fd700a5

Observation 5e32e358-db83-4460-8e6b-db9fff747a62 · outbound

This paper cites LAST: Language Model Aware Speech Tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models LAST: Language Model Aware Speech Tokenization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.271586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.271586Z digest=sha256:931383e34f691f59c97c2d7fa8e2093e65e0451d72aeeda26ff480ce03adcdf6

Observation b94787c0-2378-45a2-b9a6-351247b3a200 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.277804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.277804Z digest=sha256:b7fcbb9b5acd6c222e33a08823f3922b42e2b39768a36aade5b3b98529717698

Observation c4e6ed73-6450-4012-9043-9e684426521a · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.282895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.282895Z digest=sha256:af57ae3b329677e82a134b2fe6f9700b25a98908ab86b4b7abc82df3e4b69613

Observation 20da324c-57e1-408d-81d9-26b3b92de506 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.288620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.288620Z digest=sha256:0885cbdc6d420ba6dc6087897dffa488223786fd3383a772c17627b3164d76ae

Observation 0ccb4176-9e30-49ea-9647-bda5a5e460e9 · outbound

This paper cites CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.293507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.293507Z digest=sha256:ac182c94577563e38023953f039f2377e2d5753b1397405e7d51f23c16675688

Observation f66ab1d4-1d32-4423-afed-fe1d397bd964 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Soundstream: An end-to-end neural audio codec

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.298375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.298375Z digest=sha256:ee1dfc4e871d557a3e8ff00b6cb2a24eda1a2991bc8fa996385467ed497c65ab

Observation fe38e239-0e49-4449-b1f2-25255f4a6132 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.303523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.303523Z digest=sha256:715c996d67f0fa47dff2fe2555ca3e8ebe359fad223dd62f4a6af7498fb6d96c

Observation f1f7c4c5-f1e1-4549-af60-efd7686a3e4c · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.309151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.309151Z digest=sha256:258d052715f5cc32eac205495adda98d6dcb20f499acca49670d240fee639fdd

Observation 5b3a6cc5-aa00-47a4-b7af-0fbc3876389c · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.319277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.319277Z digest=sha256:d72c60ab81e1345bf13f6ad5d5a746d86a55c3d2ffaa4a059759d46fda04fa86

Observation be3cf178-6cc1-4782-80f5-95f27d530a18 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.324529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.324529Z digest=sha256:4412353332adc61530d68e7d0ce8087fb149e0357739c7272ca12cbf26fdc3b8

Observation 5429ead6-1b82-4c58-b02b-1dea3c8f39f5 · outbound

This paper cites @esa (Ref.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models @esa (Ref

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.330235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.330235Z digest=sha256:cc14f536ebd9c311c82e41cbe445f9aa069f6bf002eb171f610c0ad6e40a5d7a

Observation 1c77e3bb-9bf3-4d61-80c0-397382e4feb9 · outbound

This paper cites an unresolved cited work.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.335195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.335195Z digest=sha256:c11143d265617d619447e00ff72e56ac8708b8baed27f755345dec680bd76e3f

Observation db04fd29-6e89-4cdd-9c95-6c6bac061f78 · outbound

This paper cites One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.340049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.340049Z digest=sha256:7be81c71e875db2a02434e8ff54e6ec2607e4e0a10f8c8c54f75daf990e67a44

Pith citing papers

Observation f56f709b-b828-435b-b0ee-482180d02993 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.327083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:87ae38dfc1a8d099a0fc3a90958c71e5d525625ff555f15ef76fe4e84d90e23d

Observation 73a32923-2621-4c7c-9c73-e2c763ebfd53 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.643075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:9c7cfa525588b6a08c21ef35b12913f20aafbd762f9e8a0235bcac412397766d