Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:42.078895Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 11 inbound Pith citation observations for arXiv:2505.15670.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:42.078895Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:39.489566Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T00:04:22.333058Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 608b3f37-327f-4f88-835f-119c63fa7a90 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Speech, as a natural interface for human-computer interaction, is a key part of this trend
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b6a9737-094d-4609-88eb-c1b000bbc7ba · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b9f11c-8d29-43ee-ba4c-b556f2eee5dd · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model As shown in Fig
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 79f804ef-dbc8-48b1-a57b-ac24569865ea · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 99396500-d067-459a-9845-140e8de2e21e · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Training Details We implement the model with PyTorch using the NeMo Toolkit [32], and the model is trained on 32 A100 (80G) GPUs with a batch duration of 1000 sec per GPU
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4d4a84b0-866b-4c43-8599-934ee54ae2f6 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Conversation and Speech Generation Quality We first evaluate the turn-taking and speech generation quality of our model in Table 2
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6833835f-f91d-4e7d-9e7d-4ba2e93057f8 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Our data-efficient approach maintains end-to-end modeling of conversation reasoning and behaviors
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e61263e1-31a4-4cb4-b4f6-7069e922801f · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f045280f-02f7-4c23-a8ed-63968d3b255a · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Language Models are Few-Shot Learners
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1a562f-c683-4837-920a-c238882f3c59 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87bfc2ab-8c94-494f-91e6-56f030755e9d · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model GPT-4 Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3efeab3a-be21-477e-bf0f-88c8d0026d4f · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3a7013-6dce-4a0d-8559-83742fb163db · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Prompting large lan- guage models with speech recognition abilities,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a37d437a-5e04-4921-b1a6-cdd4a1f11072 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10fc18f8-d107-4c65-b84b-79018482b431 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Salm: Speech- augmented language model with in-context learning for speech recognition and translation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f8de665c-47f6-46ec-bace-b58b78ad2f12 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5de911f-2e4c-478b-ba2d-aa3d87b706db · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Chain-of-Thought Prompting for Speech Translation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe7e413-a417-44d9-9241-a2d73bb7e9e8 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Audiogpt: Understanding and generating speech, music, sound, and talking head,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a6c0b4-dc2f-43c3-9ed7-5e368dccc600 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38142bf8-5591-46c9-98fd-aaf614f8ae83 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39200706-e687-4c09-9335-2c5b907af55e · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ace2e86-a54f-4556-bce4-86c0b5df85a0 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Moshi: a speech-text foundation model for real-time dialogue
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db328e41-823a-4af7-af87-23a964b4f75d · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d69e0e4-cd4a-4954-91cd-920b36bfd365 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Language Model Can Listen While Speaking
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85dcecb1-992d-4578-8b26-c89e185417b4 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b497fc-fbcc-4410-a8e5-9be4634c3e0a · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4063eb1b-f800-41e0-9f90-47c4c9ea75bf · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5f9da1-3b89-427f-b157-5f391fb774f5 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb5ef39-3078-4bb0-824d-ab55b1f8e210 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model STT En FastConformer Hybrid Transducer- CTC Large Streaming 80ms,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d0a9ae32-ff49-4c84-ad16-77cff57fa3db · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Tinyllama: An open- source small language model,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f359ab48-b432-4563-8c82-6e8c4bdc8a3c · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Nanocodec: Towards high-quality ultra fast speech llm inference,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 48467ded-daf0-47ac-816a-8968de5ac28d · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model NeMo: a toolkit for building AI applications using Neural Modules
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7766c92-50da-4dff-adca-bdeaffa7b8f3 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c2f78b49-9ca8-4db3-baac-27a63f3685a1 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Finite Scalar Quantization: VQ-VAE Made Simple
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be44fe8-916e-4a1c-96e3-392b51f16bab · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d149fbc4-f14f-408b-bf49-7fc3feaa4627 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afeb7db7-fe2c-4f92-81b3-1c2bb52de6da · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Stanford alpaca: An instruction- following llama model,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0303236c-1072-4fcf-97c2-6f6758f0a3d4 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b269f047-b8b3-4a58-8d0d-590bb98753ca · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Everyday conversations for llms,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 81779223-e5d3-45f2-ad43-c20072f0cc4e · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff242c2-36ba-48cb-9490-3526bb759a91 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09695c8b-9ad6-4b6b-acb4-df096bd30edc · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3a9b70f2-c4a3-49f7-ae7a-391d68da4bcb · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Scaling speech technology to 1,000+ languages,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93cb78a0-cf8f-4920-a0ab-345b6f8ee6d4 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3709e300-8916-4d0d-b697-2a6974bb8e19 · outbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6a9737-094d-4609-88eb-c1b000bbc7ba · inbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997ee9af-047a-44d5-86ea-3c94a2f4b77c · inbound
Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 75cea517-e480-45b8-807f-ba283477bba1 · inbound
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 980968fc-275c-4fee-b843-018ff973d641 · inbound
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b0ce6e2a-d601-4417-9067-93248a6d57d2 · inbound
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d2eaf1-5dd8-4436-85e2-856c68393d1f · inbound
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6a58ff2-e7c1-42cb-bd98-2ed51e7dc396 · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 52cd47b6-9d6c-4c71-9cbd-5ea6107984d4 · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 23729f1c-b5df-4f87-87ea-214360879b92 · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 21092fd7-6dc1-4211-98e5-be86fc091573 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 195
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 984fd4c7-7dc4-4b3b-a0ba-379ea04b27f1 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Reference 195
Source-reported events for the cited work
Unavailable: canonical work link unavailable.