Pith. sign in

Paper Citation Record · LEDGER

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

As of 23 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 11 inbound Pith citation observations for arXiv:2505.15670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15670 v4

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:42.078895Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:39.489566Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.333058Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 608b3f37-327f-4f88-835f-119c63fa7a90 · outbound

This paper cites Speech, as a natural interface for human-computer interaction, is a key part of this trend.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Speech, as a natural interface for human-computer interaction, is a key part of this trend

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.444460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:39.341990Z digest=sha256:ca1e2ebc96d89b9a512838e5484991c0aaa97797844e298b6c9717aa4d3d2a16

Observation 0b6a9737-094d-4609-88eb-c1b000bbc7ba · outbound

This paper cites SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:39.489566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:39.489566Z digest=sha256:b5f444b992b1e92a3972cb0cb464dd972a616295c9b53886cd3159d020c118f0

Observation f1b9f11c-8d29-43ee-ba4c-b556f2eee5dd · outbound

This paper cites As shown in Fig.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model As shown in Fig

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.433740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:39.687024Z digest=sha256:8b7e536e3e038cf124a24ee9dc302752b220c2c59107cb2ff9d3335d28d967ac

Observation 79f804ef-dbc8-48b1-a57b-ac24569865ea · outbound

This paper cites an unresolved cited work.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:42.425804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:39.823483Z digest=sha256:21a0911539a5a5062f65240a13c3dabd34cefece70d57edddbe2bfc654dfbed8

Observation 99396500-d067-459a-9845-140e8de2e21e · outbound

This paper cites Training Details We implement the model with PyTorch using the NeMo Toolkit [32], and the model is trained on 32 A100 (80G) GPUs with a batch duration of 1000 sec per GPU.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Training Details We implement the model with PyTorch using the NeMo Toolkit [32], and the model is trained on 32 A100 (80G) GPUs with a batch duration of 1000 sec per GPU

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.417215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:39.932634Z digest=sha256:0f6a06b9d93302c5942f85a59b6649907360147a96d4cbf581c8b3e88921b7fa

Observation 4d4a84b0-866b-4c43-8599-934ee54ae2f6 · outbound

This paper cites Conversation and Speech Generation Quality We first evaluate the turn-taking and speech generation quality of our model in Table 2.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Conversation and Speech Generation Quality We first evaluate the turn-taking and speech generation quality of our model in Table 2

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.398893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:40.216367Z digest=sha256:62d9fa3393be6d19fd6fc8321cb1a98789514e43f5d773a1c75ec44d66a92440

Observation 6833835f-f91d-4e7d-9e7d-4ba2e93057f8 · outbound

This paper cites Our data-efficient approach maintains end-to-end modeling of conversation reasoning and behaviors.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Our data-efficient approach maintains end-to-end modeling of conversation reasoning and behaviors

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.389893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:40.284772Z digest=sha256:b9b7b637f1b2e47c67e3fb5f5c940f1c229068c10c876f048a31da3383d81e0b

Observation e61263e1-31a4-4cb4-b4f6-7069e922801f · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.710921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.710921Z digest=sha256:8d5d25137f9c0f0496a281838b8a2c4b1ab24d0af77cdb43f104c662eaac05af

Observation f045280f-02f7-4c23-a8ed-63968d3b255a · outbound

This paper cites Language Models are Few-Shot Learners.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Language Models are Few-Shot Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.323064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.323064Z digest=sha256:099b4534d0221ace36561da4c406738b07412230f8e44aa39ec323246e2f4f1f

Observation 1e1a562f-c683-4837-920a-c238882f3c59 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.358344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.358344Z digest=sha256:bcb4b9f1d46b9778d782fd90c9dfc2da26951a3266de68a5c1c6aa134ef50623

Observation 87bfc2ab-8c94-494f-91e6-56f030755e9d · outbound

This paper cites GPT-4 Technical Report.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model GPT-4 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.458268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.458268Z digest=sha256:59062979f45fcd1ef79ce8d5fff2611a513d2cba747a6b8f74b8aae8668c4dbd

Observation 3efeab3a-be21-477e-bf0f-88c8d0026d4f · outbound

This paper cites The Llama 3 Herd of Models.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.495309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.495309Z digest=sha256:591a60fc18ca135314dc06c22ee7ad16f628aba0f53e2b427258b724964d2061

Observation aa3a7013-6dce-4a0d-8559-83742fb163db · outbound

This paper cites Prompting large lan- guage models with speech recognition abilities,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Prompting large lan- guage models with speech recognition abilities,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.381098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:40.505760Z digest=sha256:deb0a8547692e46af3cf4f4741ad658a240be9d5683705167a7d453f3ae35df4

Observation a37d437a-5e04-4921-b1a6-cdd4a1f11072 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.576987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.576987Z digest=sha256:167e3951bd8ae722adc3c0b4a0a216c68843a8400063ea8af98a3bc87c3cb535

Observation 10fc18f8-d107-4c65-b84b-79018482b431 · outbound

This paper cites Salm: Speech- augmented language model with in-context learning for speech recognition and translation,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Salm: Speech- augmented language model with in-context learning for speech recognition and translation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.373175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:40.654052Z digest=sha256:d7422ac3d7295a888d1869ace1b2fc5fc5848ccd6307c316d158541c7d96fde8

Observation f8de665c-47f6-46ec-bace-b58b78ad2f12 · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.249905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.249905Z digest=sha256:0012cbb51b90c04331734f24f2a01eecf4d9037a1aadc4e6ce8e6c82e62dd1e5

Observation b5de911f-2e4c-478b-ba2d-aa3d87b706db · outbound

This paper cites Chain-of-Thought Prompting for Speech Translation.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Chain-of-Thought Prompting for Speech Translation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.786047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.786047Z digest=sha256:88f63f17b2f0f52bf7e1f068b31a6e28128693a60e43ab73b5572e53a1d27524

Observation 9fe7e413-a417-44d9-9241-a2d73bb7e9e8 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Audiogpt: Understanding and generating speech, music, sound, and talking head,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.841899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.841899Z digest=sha256:99962097423e2780952fd60fd621eb075f0b3abd55a9cb4eafee7e5105257854

Observation 01a6c0b4-dc2f-43c3-9ed7-5e368dccc600 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.897400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.897400Z digest=sha256:fe72e7faa98f3bb8c9e9a17f706fb374fa3a6cefb109ffca840be8429a06aad4

Observation 38142bf8-5591-46c9-98fd-aaf614f8ae83 · outbound

This paper cites Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.984202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.984202Z digest=sha256:218ec2c7f38b24912cb21141748c85844a996c38915716da2a7a6b7f0f313da4

Observation 39200706-e687-4c09-9335-2c5b907af55e · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.059351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.059351Z digest=sha256:c41a4525838f8d511cc3950ddd994f130675888f0ee6a2f0b786936245ded7c2

Observation 0ace2e86-a54f-4556-bce4-86c0b5df85a0 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Moshi: a speech-text foundation model for real-time dialogue

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.113033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.113033Z digest=sha256:6a4e2f58401482159113d62fc0cb2c103ee55239df3665039ef222548f595de4

Observation db328e41-823a-4af7-af87-23a964b4f75d · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.192273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.192273Z digest=sha256:f84302aba97c701e058e56014776f6525c580b815e985db89140c70efd065897

Observation 6d69e0e4-cd4a-4954-91cd-920b36bfd365 · outbound

This paper cites Language Model Can Listen While Speaking.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Language Model Can Listen While Speaking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.553345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.553345Z digest=sha256:f76113f8d5f2377193683532aca45da3f40f0e49b19cd5db33cd36a2057b561b

Observation 85dcecb1-992d-4578-8b26-c89e185417b4 · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.290160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.290160Z digest=sha256:a950b69f85b321d054a600f49c05818c73c9c5360855875d09325a93a34f86f7

Observation 33b497fc-fbcc-4410-a8e5-9be4634c3e0a · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.303521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.303521Z digest=sha256:b0e26bc551868419c2f3b668d6847358b9dcdc67c1abf8a2c873e2573fe055c1

Observation 4063eb1b-f800-41e0-9f90-47c4c9ea75bf · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.330420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.330420Z digest=sha256:fe80276db18de74b714ae6c0c5e36b55d0f03dfddc41aeb9f5ecf3b6426cb233

Observation 3d5f9da1-3b89-427f-b157-5f391fb774f5 · outbound

This paper cites Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.384675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.384675Z digest=sha256:ff7d491178650915f5b112196e81e6369139277ca3a3b2c71f220bf44ca76e31

Observation 0eb5ef39-3078-4bb0-824d-ab55b1f8e210 · outbound

This paper cites STT En FastConformer Hybrid Transducer- CTC Large Streaming 80ms,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model STT En FastConformer Hybrid Transducer- CTC Large Streaming 80ms,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.360642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:41.446726Z digest=sha256:d83ba1f5a09a77dd9cded9d9f9e9147eb4dcf4425223833c53a520ecef396543

Observation d0a9ae32-ff49-4c84-ad16-77cff57fa3db · outbound

This paper cites Tinyllama: An open- source small language model,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Tinyllama: An open- source small language model,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.352226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:41.480620Z digest=sha256:7cbf88782b41f21518b40f1d5c0e130dd0dd259e5b285322b70b798a5504abcd

Observation f359ab48-b432-4563-8c82-6e8c4bdc8a3c · outbound

This paper cites Nanocodec: Towards high-quality ultra fast speech llm inference,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Nanocodec: Towards high-quality ultra fast speech llm inference,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.343551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:41.515723Z digest=sha256:0041b979322716c5902f8917bbbd04bc8532835d40bae8c857a86ea816b0063c

Observation 48467ded-daf0-47ac-816a-8968de5ac28d · outbound

This paper cites NeMo: a toolkit for building AI applications using Neural Modules.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model NeMo: a toolkit for building AI applications using Neural Modules

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.858736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.858736Z digest=sha256:d6986982061b841588ef9b8095fd5c15e8e56c0d5ceb56dac0283214c4977b14

Observation b7766c92-50da-4dff-adca-bdeaffa7b8f3 · outbound

This paper cites an unresolved cited work.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:42.407141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:40.071890Z digest=sha256:491c15e14f5d155c5eef9f1fec3b61453d8022c6e0f5b0a96c70ba2a58f9cf70

Observation c2f78b49-9ca8-4db3-baac-27a63f3685a1 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Finite Scalar Quantization: VQ-VAE Made Simple

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.585656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.585656Z digest=sha256:a659c3f91e3003e7d89cbbc704445aa6ac45ce4d54470b1d4aeb230e247aa900

Observation 8be44fe8-916e-4a1c-96e3-392b51f16bab · outbound

This paper cites Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.620262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.620262Z digest=sha256:b0dd6083871b3e0c0cb00505430047d0c271c5d61cc49778cb1b335368bdd5cc

Observation d149fbc4-f14f-408b-bf49-7fc3feaa4627 · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.665957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.665957Z digest=sha256:142c6d8427443340af2383ea2d5effa0ade07ad6147fdbbbaaa702998ed0ccfa

Observation afeb7db7-fe2c-4f92-81b3-1c2bb52de6da · outbound

This paper cites Stanford alpaca: An instruction- following llama model,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Stanford alpaca: An instruction- following llama model,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.335636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:41.701071Z digest=sha256:32eacfcc27c0b0020ad924546972067dd9eef2851ccd8476f69ce5a043a17999

Observation 0303236c-1072-4fcf-97c2-6f6758f0a3d4 · outbound

This paper cites Instruction Data Generation and Unsupervised Adaptation for Speech Language Models.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Instruction Data Generation and Unsupervised Adaptation for Speech Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.737032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.737032Z digest=sha256:1fcc317dc34220ef6fd9c234e5644c059c28a3677c48cae9e05cbec37024432a

Observation b269f047-b8b3-4a58-8d0d-590bb98753ca · outbound

This paper cites Everyday conversations for llms,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Everyday conversations for llms,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.326808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:41.770206Z digest=sha256:6217e73c37557b66c599015943a7b04fe17096752d5873a75defa6aa6231fe02

Observation 81779223-e5d3-45f2-ad43-c20072f0cc4e · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.815909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.815909Z digest=sha256:3e414eef987bf9050235e9df68376b5d88620a9a9d266a90a28550d3b096c0ce

Observation 4ff242c2-36ba-48cb-9490-3526bb759a91 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.893922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.893922Z digest=sha256:29c61b2e95bdcd9fc799d9f742f215c82503cabf569483e5eb0161d62162b070

Observation 09695c8b-9ad6-4b6b-acb4-df096bd30edc · outbound

This paper cites Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:42.317772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:41.945864Z digest=sha256:28dd3be30c6d3c454a62554c19012ff8aa09e231acdc9e6e75f4064941332151

Observation 3a9b70f2-c4a3-49f7-ae7a-391d68da4bcb · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Scaling speech technology to 1,000+ languages,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.996729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.996729Z digest=sha256:b6aeb97cace60e054e00b389f1840514ded40dd58590b369853a12b9221bf5f9

Observation 93cb78a0-cf8f-4920-a0ab-345b6f8ee6d4 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.069009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.069009Z digest=sha256:9fef0bfcd750653cf713750ef57b6ec7839c6dfcf96c67526d341dd683fdc089

Observation 3709e300-8916-4d0d-b697-2a6974bb8e19 · outbound

This paper cites ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.078895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.078895Z digest=sha256:5a622a616f0964610acd9762caf5f628baefed80dac968e6cc181b1861f9c4ea

Pith citing papers

Observation 0b6a9737-094d-4609-88eb-c1b000bbc7ba · inbound

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model cites this paper.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:39.489566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:39.489566Z digest=sha256:b5f444b992b1e92a3972cb0cb464dd972a616295c9b53886cd3159d020c118f0

Observation 997ee9af-047a-44d5-86ea-3c94a2f4b77c · inbound

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models cites this paper.

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.609029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T07:43:23.913399Z digest=sha256:f7a7c4226f0e883d943c3d23b53afd27e48f70a960f6976017e9e6f5337d6458

Observation 75cea517-e480-45b8-807f-ba283477bba1 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:35:18.353408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T08:34:56.898815Z digest=sha256:93ee1be1f4981a694045bcdb57ddb24608d95493b68f36ec40a151cf8fd28d0d

Observation 980968fc-275c-4fee-b843-018ff973d641 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:44:07.766227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:43:27.176535Z digest=sha256:5641667b6b877eb72ab2dfa38c24fbe94e0943c8b789030151f728afa9c97635

Observation b0ce6e2a-d601-4417-9067-93248a6d57d2 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T22:57:11.059962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:57:11.059962Z digest=sha256:3808dc52a100fa479d288f1c0679e37fd86648c17d6b6935381b3a01590dc917

Observation 38d2eaf1-5dd8-4436-85e2-856c68393d1f · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.873747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:0ec9f520e0df2e0503112d27997ca5ca4b4d1274d049aea98200e6eb3a0038b4

Observation a6a58ff2-e7c1-42cb-bd98-2ed51e7dc396 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.679040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T00:13:38.523397Z digest=sha256:366a90bbe75421864c9768b2b5f94f234f6745d0b0b9bf2461f52370e2499c4f

Observation 52cd47b6-9d6c-4c71-9cbd-5ea6107984d4 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.969399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T05:27:24.078053Z digest=sha256:29641c78423723dae0436f9dbd3dd1e12177f8f59fb5c670aac6a163d0fa27f4

Observation 23729f1c-b5df-4f87-87ea-214360879b92 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.305137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T09:43:09.347118Z digest=sha256:817cbc3119ec54898a90f0abbf787368a78b378679b5a3b04ddf51605aafc2e9

Observation 21092fd7-6dc1-4211-98e5-be86fc091573 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 195

Resolution
verified exact
local_arxiv, observed 2026-07-08T00:04:22.334602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:59aa7d14ff0281dbcbed3ea8b63132ff07e0e38cda84bfbe60cc05f3d9f9fafd

Observation 984fd4c7-7dc4-4b3b-a0ba-379ea04b27f1 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Reference 195

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:0c805254d1ea8981fd6145686f3b93c91920b58a3a222eb8f9f584ba34026793