Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T07:07:57.800866Z
Paper Citation Record · LEDGER
As of 23 July 2026, this Paper Citation Record lists 41 of 41 outbound references and 58 inbound Pith citation observations for arXiv:2306.12925.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T07:07:57.800866Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T20:46:17.286576Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T07:56:57.816571Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f6eec113-0800-4141-84fb-7a9173153aed · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen MusicLM: Generating Music From Text
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2cc3e3ba-e23a-41a7-95ba-2a4f6cd02634 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen PaLM 2 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3a23f3e7-60a8-4889-866f-ae7823c9377d · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen ISBN 979-10-95546-34-4
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 46595cb1-d75f-46f7-8141-624fc939a5c5 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen mSLAM: Massively multilingual joint pre-training for speech and text
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a0d0240c-b785-4828-89a6-b9447878a1c0 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Barrault, O
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8fe03784-b822-4e43-aacf-6408d5176461 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7561c7d2-fd9b-4451-92c2-c521f32e5a6e · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen org/2020.wmt-1.1
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7bf0ae9d-5eb7-4b4a-a19d-c0a7b8943226 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6ddc9af9-8a1a-45c5-98ae-ab1b778e9d37 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation eee3fbf9-fade-4b92-bf9b-2d71fda20473 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3803036b-ccec-42d2-b2ea-823ae98ff3a1 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen AudioLM: a Language Modeling Approach to Audio Generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 42dd9788-f40a-4fc5-ab55-c6d196393a9b · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen SoundStorm: Efficient Parallel Audio Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 81cb6522-d1b6-4dad-bc16-86fd86254f03 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4db85b1f-9af3-47c2-8dc6-eb18a317ecb4 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 23b77fbb-f4af-495a-adc8-a8f96f9b7c33 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen MAESTRO: Matched Speech Text Representations through Modality Matching
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation dc54923b-933c-406d-87c1-c16c58f5a204 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen PaLM: Scaling Language Modeling with Pathways
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5c1e4ed4-24f7-4478-bd5a-efb10731b9e1 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Conneau, M
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 461a446a-d472-4a72-9ff7-6c52f7f8362b · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen High Fidelity Neural Audio Compression
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 005414fe-954b-43e6-9562-48fbc81b2a72 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen High Fidelity Neural Audio Compression
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b1679d38-3718-44eb-abc5-c42791bf174d · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen SingSong: Generating musical accompaniments from singing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9afb53b4-24ef-4935-8638-e76401a69ac4 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen SingSong: Generating musical accompaniments from singing
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5a7b3691-3bf3-419e-9e4f-249fe9612f82 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d1c8f20b-cb0a-4f0a-b03c-fd8f2e450c29 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Textually Pretrained Speech Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1081a94c-384d-41f8-b430-fb48a3de3db6 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9536e6d7-2c9a-4829-bffd-cf841a63969b · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 18efb642-dc83-4d3b-b56f-8a5f624116de · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ead98b7d-9584-4ad0-a34c-abdfe64ff9ed · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen AudioGen: Textually Guided Audio Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9a1c68fa-3bca-4550-bce5-b4e39b77681c · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen AudioGen: Textually Guided Audio Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation bb76f95e-ce07-4342-bcc8-7cc4274a8ae2 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Textless Speech-to-Speech Translation on Real Data
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4b0fa63a-8a40-4832-8522-6c9979623af7 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Direct Simultaneous Speech-to-Speech Translation with Variational Monotonic Multihead Attention
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f828482a-e636-4e2c-8a23-f613b6e71548 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Representation Learning with Contrastive Predictive Coding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation fd363f26-c4b4-412d-af61-829595318744 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen MLS: A Large-Scale Multilingual Dataset for Speech Research
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 469aeb7a-379b-40b0-aad3-4a1b44334166 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9f8e0cfe-3407-4564-86fa-c8007dc34f89 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9c409c77-d2f6-46a8-bb54-13ea20ac87b2 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Robust Speech Recognition via Large-Scale Weak Supervision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation dbf0ed76-bf9a-4c19-9fc2-6df06ffb022e · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 0be3195c-1df0-4d4d-af57-6e9552c6d881 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen CoVoST 2 and Massively Multilingual Speech-to-Text Translation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 40f490d2-0806-4da5-bdb6-2be2199dab83 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 15ec8616-8e22-470c-97aa-bef1f122e1ac · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6ac8d2d5-b915-4ac9-bc4a-e8f79f515d34 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5d1ca95b-e901-43c1-ba66-8378f6cc8a70 · outbound
AudioPaLM: A Large Language Model That Can Speak and Listen Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6d5d19a8-ec8e-4dcf-aaa0-1e930b71436d · inbound
A Survey on Multimodal Large Language Models AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9b502f34-cb70-408c-bd87-4055a2a6e8ac · inbound
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 216
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c4bb07a3-e80a-4bcb-b492-d8e23daf3585 · inbound
SALMONN: Towards Generic Hearing Abilities for Large Language Models AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9bae3512-608d-42df-8cee-aca6de0ea124 · inbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d49019ca-98ee-4907-bd2b-a437cbe31600 · inbound
DASB - Discrete Audio and Speech Benchmark AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 845d86b3-8597-459f-b362-021156237d04 · inbound
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f452b515-2a6f-42d0-8791-4a26ba2947de · inbound
Moshi: a speech-text foundation model for real-time dialogue AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 320f1ed4-37b4-4164-b853-dc288ec0e6fa · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e14ce1df-f34b-45f9-8f90-b6a18f7055aa · inbound
Step-Audio 2 Technical Report AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 56de3074-c47f-4837-8891-9d7bbdc6be87 · inbound
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3401f375-34b5-430d-99ab-5a502de98a97 · inbound
Enhancing Speech Large Language Models through Reinforced Behavior Alignment AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 49e37056-2567-47ce-9538-73a740eb7c55 · inbound
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting? AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 10c8f76f-1eec-47fe-822c-b669a9b6b5a4 · inbound
Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d27a1e66-cd50-43c1-99b8-4b0e0fda3db7 · inbound
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1c982429-0c85-4830-9ecd-da24a5a83063 · inbound
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8e72485b-2cc6-4b9a-b5fa-6a56e2673251 · inbound
Generative AI in Signal Processing Education: An Audio Foundation Model Based Approach AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b65cae82-50fa-4474-8a71-11e63c0c172f · inbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 23f4929f-2385-44b7-8bab-2fc0103db38e · inbound
LLMs and Speech: Integration vs. Combination AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7db8a1d4-4cd8-44ac-88a0-9bfad8030621 · inbound
LLMs and Speech: Integration vs. Combination AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4362face-5bed-465d-b259-41e07951cfe8 · inbound
GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 97ec44a6-bc12-4425-babb-b6a8e751cd52 · inbound
Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f1fc358f-0d6b-4632-8dcb-cf87c439a3de · inbound
ViLL-E: Video LLM Embeddings for Retrieval AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e1475385-ea6f-4c0b-90d8-222a94cae204 · inbound
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 697e3bb3-066a-4ab8-86d3-3df871aa772d · inbound
Beyond Feature Fusion: Contextual Bayesian PEFT for Multimodal Uncertainty Estimation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 20bc72ea-b452-4915-bea1-fa8d40bfdb1b · inbound
ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7086e60b-dd63-4f3e-bbd1-055b11621a4e · inbound
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 44c0cbd6-526b-415d-a0a5-de2ead981b81 · inbound
Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 466c03b9-9196-45a1-b0a0-dd2b53f4ce99 · inbound
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f41b4629-0450-4767-8953-2baa927df29d · inbound
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation aaa2c169-7f71-48d4-872e-af410ee72835 · inbound
Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation abf47f2f-ba85-43ad-9797-4b86d9ab3211 · inbound
Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d48dc5b6-11f7-4499-896c-35402376e954 · inbound
AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation aa595d77-54f3-479c-9406-a96890a868ae · inbound
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d6f59265-9359-4906-b01c-3441777c4884 · inbound
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c07f6c49-bd15-4d72-82cf-5fcb0c2c5ab0 · inbound
Direct Translation between Sign Languages AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a0b94de7-b2d8-424e-a57b-dac836006370 · inbound
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6bf2e881-a9ad-4f4c-b240-68ab16f04ba1 · inbound
Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models? AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 26a40d52-1137-4b4b-9451-cea9a95f1966 · inbound
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b6454bdb-ef7b-4124-b450-3705762939a6 · inbound
Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 81d14c45-8060-4dad-901d-2ada3f13f65e · inbound
Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f5b65a5e-84a4-4f2a-9e20-ebb402c64e0e · inbound
Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5d036bd6-9f2a-47ad-ba37-db9e90e10ce6 · inbound
ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f81c375d-f0ab-4f3e-8be3-771619bdadab · inbound
Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 262
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 56d51efa-b234-4d4f-a7fd-1306ac4557d8 · inbound
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5d760ff3-1def-44b4-974b-aa41bc8ba86d · inbound
PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6861a717-7b75-4f71-b0af-b0f9deb946a7 · inbound
Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f9c5f736-22d6-4750-9c1d-33408eed34dd · inbound
Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 86901faf-a1a2-46f6-b70e-cc3fe8ad8977 · inbound
AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6113f7ee-e4b1-4231-a088-fb8bf09941d7 · inbound
Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4bc40f73-88d8-4062-9086-1ad926b14211 · inbound
Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e0a2e54e-5510-462d-b657-e2650a622d66 · inbound
How to Leverage Synthetic Speech for LLM-Based ASR Systems? AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e6e48d35-7858-4ab0-a4bd-b50b870ffa7d · inbound
How to Leverage Synthetic Speech for LLM-Based ASR Systems? AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b66131a4-129b-4253-b2d9-870cc4b0d8e3 · inbound
FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7d620171-28d8-4eed-974b-bfb03e61101f · inbound
Neural Audio Codec with Adjustable Token Temporal Resolution Using Sampling-Frequency-Independent Convolutional Layers AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 393dd19c-04e8-4b2a-a6b2-35c522eb076d · inbound
NAVER LABS Europe Submission to the Instruction-following 2026 Short Track AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 74155a37-6d9e-4e87-904f-4ae381eba66b · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 238
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e4c9bea6-6239-4ca3-82ad-920ffef0b665 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 238
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a95b65-001a-4544-bf80-7d7c79d6f798 · inbound
When Synthetic Speech Is All You Have: Better Call GRPO AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.