Pith. sign in

Paper Citation Record · LEDGER

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

As of 20 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 10 inbound Pith citation observations for arXiv:2505.17060.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17060 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:49:58.269093Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:34:50.786808Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:28:39.133843Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f8bc5e1-3415-4bd7-ad9b-2b64f7f5fdd2 · outbound

This paper cites GPT-4 Technical Report.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.085162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.085162Z digest=sha256:bc1517e1a0e477239187de11aa471a6e04703a9d5bbe150711180596aa97ec94

Observation 86422141-0234-4b2d-b414-a83dce500320 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.088931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.088931Z digest=sha256:35a0c9e797604187ad4b46acb73cd9d75b6433d31eca25d79a3d143525a4b4a6

Observation 17f35836-3acd-4093-bd65-55b3483352ec · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.092125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.092125Z digest=sha256:b4d3e623a58565b8c7e9b2dee606c0f469c730be987e7792735ecb7bc5c444e9

Observation 4c366b08-f12e-45a1-a39a-b2b647574e79 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation SALMONN: Towards generic hearing abilities for large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.651721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.095582Z digest=sha256:f5fa128e53c7eea29d9ef0e6f811d983283c931539e4bda4ca5ca1faaf771856

Observation 54817caf-0997-4b07-9e7e-45c544ae5cc0 · outbound

This paper cites Listen, think, and understand.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Listen, think, and understand

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.643479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.098927Z digest=sha256:231ed26bde2b8ca410a681b247c419313005adcf4229331df83c4a69d4e20839

Observation e003506d-3d82-4363-8356-4466907ab0c8 · outbound

This paper cites Qwen2-Audio Technical Report.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.102213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.102213Z digest=sha256:315a6293bb66dc4282897d270626fefd29387e4571875add1d301e8d84e627c3

Observation 93bec0d0-34bb-4c52-845a-a2456af118db · outbound

This paper cites Boosting Large Language Model for Speech Synthesis: An Empirical Study.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Boosting Large Language Model for Speech Synthesis: An Empirical Study

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.106280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.106280Z digest=sha256:b490b180a1025cf157682f4db6fc55424ce6a0d1cffde9ec8e728e8e6a36f205

Observation 675fa110-ab79-4cd1-9ef2-3fcff72d1255 · outbound

This paper cites The Llama 3 Herd of Models.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.110289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.110289Z digest=sha256:8134bc56fe4cc18481ac50450cdb8a4a159568f888df49db7766aa63f8b43ad5

Observation 4c1f2931-6378-4d68-9f1a-5eb2208e23c5 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.118039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.118039Z digest=sha256:f381a3611ee8cd6562f7894b1efac7540fab473543ec4a8123f3539625817f39

Observation 0d9e9fae-f321-4dbe-b423-16acc5ea36b7 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Moshi: a speech-text foundation model for real-time dialogue

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.121520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.121520Z digest=sha256:93b261325e6b82370a7cd6e1be3c10387c661372f54e7677d4f491652306f592

Observation 56099895-2d65-4127-bd92-c82389113659 · outbound

This paper cites LLaMA-Omni: Seamless speech interaction with large language models.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation LLaMA-Omni: Seamless speech interaction with large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.629655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.125030Z digest=sha256:a45fd9d52ae1f6d074c66b376ddead7e6445bc4d6dee069437b7d46b46d52703

Observation f90d1971-022b-41b1-9452-738d00b77af7 · outbound

This paper cites Spirit-LM: Interleaved spoken and written language model.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Spirit-LM: Interleaved spoken and written language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.619872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.128306Z digest=sha256:076784e98033e546437fa362842677327d72ac5c0a06ab2077dcd25135de10bd

Observation 231eaac7-9fc4-4c13-ba36-66127ba8ffe7 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.131650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.131650Z digest=sha256:7b3a4275c47ce2bd4f3c7d10cacf758464a04bce43d766e0fb3ebd5111f6f09c

Observation 739b44ef-b6a2-489e-bdfd-486f9aa8cf9a · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.134642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.134642Z digest=sha256:df9e560e3b0821d86af053e2778e4f8661239715ed075ac2047cde7837d77a07

Observation ecb017c2-ac14-4fc6-a3d0-b50eb6ea12fe · outbound

This paper cites Beyond turn-based interfaces: Syn- chronous LLMs as full-duplex dialogue agents.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Beyond turn-based interfaces: Syn- chronous LLMs as full-duplex dialogue agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.137540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.137540Z digest=sha256:3dec566f7c022f6a60d3f029a8405ac3a1232f1d1bcc846f3bc23b1b24e99583

Observation 2837a4c3-fcda-4c60-b9a3-ff1ba5d66d08 · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.140836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.140836Z digest=sha256:22546a82b07e0e032ea66061d3adb3a185a6c01d6a65f04c139d4b7619bd83eb

Observation 50a42bbd-10ad-41f3-84e1-02ba9cc7c3e9 · outbound

This paper cites Talking turns: Benchmarking audio foundation models on turn-taking dynamics.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Talking turns: Benchmarking audio foundation models on turn-taking dynamics

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.607870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.143685Z digest=sha256:03e699c0f0ba9631a1f55c187d3156a14558d916d998456ee39549a73843aede

Observation 583367b8-f255-416f-b385-d2ac3dd7fbd0 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.147040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.147040Z digest=sha256:4c0706343419ab37c316f6ffab984e9578d92548deb6cae126d4613c826caa50

Observation 79305b73-1c7a-448a-9778-1edf396083e3 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.149680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.149680Z digest=sha256:ac115366971b803ccef6094a3a459b846ed3df1711c0fcb6cc35461ce1fca3d3

Observation 6e30344c-e13b-49bf-a9ed-1599b06c7eeb · outbound

This paper cites Kimi-Audio Technical Report.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Kimi-Audio Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.152800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.152800Z digest=sha256:8f46723bbdd8cff16c06047fceccb6783459c771a6225820b999d5409cae4077

Observation c8735e97-e5a9-4209-81d9-84a83e26138c · outbound

This paper cites Qwen2.5-Omni Technical Report.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Qwen2.5-Omni Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.156421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.156421Z digest=sha256:4bf5b3bd5c384cb87c7ab8cb7faee17c12b493819814445cec5aaea622b4667f

Observation 05e232ff-1b66-454e-a2f0-5cb41330643b · outbound

This paper cites POMDP-based statistical spoken dialog systems: A review.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation POMDP-based statistical spoken dialog systems: A review

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.159734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.159734Z digest=sha256:9c73e764b5467612236ce035fd4510066ce0422225a3905f1faf60d8b64a55a0

Observation eed314f8-dc09-4254-af30-586cfc73b979 · outbound

This paper cites A neural conversational model.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation A neural conversational model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.597565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.162275Z digest=sha256:7c05d9029aae80b47bfcb83d744bdc8760e5c5ec2454489c7eddeceb0f03a920

Observation 5fd6d634-3dea-4fdb-91f6-0ac4fc15259b · outbound

This paper cites Building end-to-end dialogue systems using generative hierarchical neural network models.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Building end-to-end dialogue systems using generative hierarchical neural network models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.590040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.166005Z digest=sha256:88648e30044f28606a24533ef5d167edc92ef5ad494948bf3180a1b3b28573df

Observation b3310856-a2ff-4889-8240-7e2a8336ff69 · outbound

This paper cites MiniCPM-O 2.6: A GPT-4o level MLLM for vision, speech, and mul- timodal live streaming on your phone.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation MiniCPM-O 2.6: A GPT-4o level MLLM for vision, speech, and mul- timodal live streaming on your phone

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.581740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.169401Z digest=sha256:7ad95922414a9451a23c04fd977237cd0b8190a42c5d864ebc37f96ee821f5bb

Observation 78287fd3-c86c-448a-a5ce-0c5c0a266b32 · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.172025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.172025Z digest=sha256:6fff680de363e42d03c3b0545d063eb6fea9b740b12a3ae8bc77c0f8fb991c08

Observation 34b70932-5755-41b0-90f4-6ac9cbb1fd62 · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.175649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.175649Z digest=sha256:92e0355dc5921b5bec7f0401f549d7acfb0cebc662b66eac2ffba7ad58b32462

Observation f3b591c2-11bb-4ae7-add7-87c52aef8f42 · outbound

This paper cites Generative spoken dialogue language modeling.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Generative spoken dialogue language modeling

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.574275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.179249Z digest=sha256:325e3425ba8c629ed8454caff43b0728761cc399325331ba976e6cd46a9afc67

Observation 67eec411-3f40-4ed3-ac22-b332e798afe4 · outbound

This paper cites Optimizing expected word error rate via sampling for speech recognition.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Optimizing expected word error rate via sampling for speech recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.568098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.181849Z digest=sha256:1f1c98b6f89832c76898205df47490301765ed488bf970cb63d38da98e4abdb7

Observation 67ef88ab-2bbc-4dce-ace0-fda76e1e0749 · outbound

This paper cites Variable Frame Rate Acoustic Models Using Minimum Error Reinforcement Learning.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Variable Frame Rate Acoustic Models Using Minimum Error Reinforcement Learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.561207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.185032Z digest=sha256:dea0721d020ea74ea6c8d79dc27549ee08646dcee67a267cf6bdbc5998491167

Observation aada64ff-0e23-4d7e-8287-9a4d7a9d7ece · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.188741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.188741Z digest=sha256:d27e1c2578a2d63dcbd0e2679dcef5a6802ca99180bef0b66b23c3f02d0b228f

Observation caaa36f8-9ae8-4c79-affc-d61d5b1f7315 · outbound

This paper cites Speech recognition with LLMs adapted to disordered speech using reinforcement learning.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Speech recognition with LLMs adapted to disordered speech using reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.553550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.191471Z digest=sha256:461b87eb6c17601a6c52f60ded464263cef2f084815280f92bcbf7876caf5e17

Observation c1a834b9-609e-4429-825f-184fcc3c9d26 · outbound

This paper cites Reinforcement learning for spoken dialogue systems.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Reinforcement learning for spoken dialogue systems

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.547112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.194107Z digest=sha256:66effa3f6a0d51d71f1384fb3c3dcbddeb710f88039ada5f5ff4dd69906c13ff

Observation f0a7409e-0094-4375-aae8-b56a72ed629c · outbound

This paper cites Automatic learning of dialogue strategy using dialogue simulation and reinforcement learning.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Automatic learning of dialogue strategy using dialogue simulation and reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.540537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.196626Z digest=sha256:b92e096a9ced0ea3b12c3b336f72295dc22d031d1ceada331179b0f327c1f3c8

Observation 7b71ea8f-e9e3-43c9-9f18-fbcedaa11903 · outbound

This paper cites SpeechAlign: Aligning speech generation to human preferences.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation SpeechAlign: Aligning speech generation to human preferences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.531123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.199241Z digest=sha256:13a1c74b8c5b0ba23c9d454e6c79bf65b2dd9ce72136525da662ebdc8b7015e6

Observation 5d4cbd6d-ae50-471b-9d80-3530c4c2bf08 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.202244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.202244Z digest=sha256:3b69e1f2f54300e5f2d6ed2a74060d6c7c13bb351ea7883b8a11dbd0797fb37d

Observation c2a291e0-faca-4c74-ab4e-794366f69776 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.206116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.206116Z digest=sha256:6dfa32df2ad012c047ab5d2976c95eeec269bb67a3313c21991b75ff40a71f3e

Observation 4f063694-7994-4176-b213-db3cbe1eb8e7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.209597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.209597Z digest=sha256:6f18bf875071c41d3d3dc2cd78159426e9cd27e404db9412c67e91e513ff9e09

Observation 6accb4ed-1745-4af5-9312-ee5eed170f28 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Direct preference optimization: Your language model is secretly a reward model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.523466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.212886Z digest=sha256:6fb94e54dfe4de2702bc5c68a49d8f582f867a81747fa682a8ab231d6e5102b7

Observation e606dcc9-01cd-4edd-9e0c-fe8c7aff0756 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.216006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.216006Z digest=sha256:9befa31df8077d2ed0d14cfef2043d850818de3910943684d9cbb3d0b5137f34

Observation 415ae299-5a84-4b92-b7b1-b46549798e4b · outbound

This paper cites MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:49:58.306059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.218894Z digest=sha256:89c82f9d5a1f72ce32563f3deb476e4ce0111df7e6599a54f7d1706b563b2a97

Observation 9db483e3-264e-4ae3-a332-59532e93ef57 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Robust speech recognition via large-scale weak supervision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.516339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.221621Z digest=sha256:45ff798f496b08373b320f028e1ee8a3ef5b34a75824d6e8be59c00fd909bae1

Observation b54a9c58-8550-4ae9-9d3e-1537a16cf97a · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Mamba: Linear-time sequence modeling with selective state spaces

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.508291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.225462Z digest=sha256:8191931098b611496cbd12f69c9cca3edd17083ea8f329913f0d8ebff9078d04

Observation 474c5e20-6729-43f7-b999-8f2a2e8b3504 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.227998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.227998Z digest=sha256:3898a29a8cef484a0fd8fc829d844bc1583e0ba871e89ddb7af1344b1b5e266e

Observation 9640781a-ed15-4829-b666-ec26ce945897 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation LoRA: Low-rank adaptation of large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.499882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.230256Z digest=sha256:7cd406fc872c7badc2aaefd818ee76a579bdf0c74869dfe138369cb81b3a73ce

Observation 75d5e047-79e7-4841-a7de-d39d0506e9f1 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Librispeech: An ASR corpus based on public domain audio books

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.492920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.232547Z digest=sha256:3b694d5b584e28ee8a40a6202d08617cb0222d2ecdc7f2d219d8d78fb7499b2e

Observation 6cbb1e05-cec9-44c9-bfc4-3277c8cf121c · outbound

This paper cites GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.485609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.235198Z digest=sha256:713d4082cf9ad1193658b2a5342841c82b767fd41694df452d59391221c6be1e

Observation 8dd047a0-5d66-4127-807b-769a79bd7b03 · outbound

This paper cites Stanford Alpaca: An Instruction-following LLaMA model.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Stanford Alpaca: An Instruction-following LLaMA model

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.477982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.237542Z digest=sha256:d9ef3bc761110a63cc6043c4d729ecb207a6565f987cb57737acc5d9f72296c6

Observation 736be314-e242-4179-939c-d826a178bb82 · outbound

This paper cites Semantic parsing on freebase from question-answer pairs.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Semantic parsing on freebase from question-answer pairs

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.468952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.240895Z digest=sha256:70c9aa06e7b80bf0af024e0144649ad3cd3f09d6905f041422d3b53046d80ccb

Observation 38231fa3-dcfd-4dad-9ac0-ae611a2456b7 · outbound

This paper cites TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.459946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.244083Z digest=sha256:722b8447a76e0bfb20128fdba6a4b867cf817b8d0ce9a7bf2a18da34c2e53e53

Observation fce24798-0a6d-468b-a8aa-0256a9f0b58a · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.452790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.247880Z digest=sha256:183f4c1f0b1721c3518a2041e712cb147f2efb0b9c5507c4598c1cc1fea106e2

Observation e0b9b03d-c329-4c72-8b7b-9198e0d20203 · outbound

This paper cites Natural questions: a bench- mark for question answering research.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Natural questions: a bench- mark for question answering research

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.443987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.250802Z digest=sha256:6c1df7ab0d0789923a18e0704c87864be452cc168bb20029299f2dfd42c23198

Observation 5e189fb4-7b27-416c-932d-4bf450e13f64 · outbound

This paper cites SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.254721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.254721Z digest=sha256:245eeb3128ac89105dbd496d9940367e2b3a9cb5bc3467f8962ec1789195bc2a

Observation aaa11f80-04e7-4137-8e45-5b6879108a55 · outbound

This paper cites Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:58.436845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.258540Z digest=sha256:c8780f36c2fbb995f6c3a9e5124b992b999717b7b2b4fa4a1d1396b647d25e3f

Observation 8d4e848b-2fbe-42f9-85f3-5c78b1ff0a44 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.261672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.261672Z digest=sha256:79ffe1b700973a287008d4012ddf95bb294b42eaf9a5ca3790f87b79340fd5f9

Observation 8f4e6452-14da-4f99-95f1-7285e8ca33da · outbound

This paper cites Libriheavy: A 50,000 hours ASR corpus with punctuation casing and context.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Libriheavy: A 50,000 hours ASR corpus with punctuation casing and context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.265684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.265684Z digest=sha256:cf5035581b647f8ca07e1b8af4e52d6792ba0d6f2f542ad2b54735f109d0bd72

Observation 40d9929c-44ad-4226-940a-9e19eb151933 · outbound

This paper cites Decoupled Weight Decay Regularization.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation Decoupled Weight Decay Regularization

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:49:58.425949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:49:58.269093Z digest=sha256:1e1a0ffb3ea5a495481f925b933b24c56dc8d0bf61431fc36110f74d04fcd028

Pith citing papers

Observation 08b72cb5-3750-498e-8192-f7b162fd8fd1 · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.704744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:bec1a5c67bbc70bcb933005cf914be992750d10bc0992d506f7e97b32c09337c

Observation 88a0d817-2b58-4e84-83b2-ca92c957c846 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:35:18.379168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T08:34:56.898815Z digest=sha256:16cd9db5480ded85a9ada98019d7ec080adc73d253ddbb5d52ae72fc030a26c7

Observation ea35f16f-6991-485c-8ad9-2ad830756197 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:44:07.707055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T10:43:27.176535Z digest=sha256:75ff9b7b84cb0c624d39fc0bde23282ba744105b8ae5a29896103e2dd939823f

Observation 4d459a23-9ee9-43c1-aa4c-a75a110836bf · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T22:57:11.059962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:57:11.059962Z digest=sha256:43dd3e6ba60b30b3612eaeba86f4843fa5214f733aef1ce839b737d494485060

Observation 2360c1d3-009a-48a7-ac00-1e304b47a8ec · inbound

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue cites this paper.

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:19.100069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:08:55.753359Z digest=sha256:b49c1df37a746dd2f305696a100bd462b16ecfac5147266f4b08e3253310b9b4

Observation 0001a4e4-39f9-4960-b04d-3ef97dbb91cb · inbound

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding cites this paper.

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 6

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T18:04:58.380348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T17:56:38.345335Z digest=sha256:e65a1f2c2a949fa656d563db19771088a5e340985ab258f13af5fa4ab22d0265

Observation 030fe35c-2674-446b-b5c0-b3fc2ffab60a · inbound

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction cites this paper.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.936721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:2cc89eebd8d3b377e526221916dd40699bc9803e4e175bcddba82a65cfa6277b

Observation ef48b538-f88a-465d-aa40-f5f40181b7a2 · inbound

Endpoint Anticipation for Low-Latency Spoken Dialogue cites this paper.

Endpoint Anticipation for Low-Latency Spoken Dialogue SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:39.135567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T05:37:17.684884Z digest=sha256:737ccaf4db41426bf9073cc69a38eaa59eb07fe71d6f8f9eae622e9bceb06f13

Observation e9a2a03c-406b-4b50-995a-adaea41f933b · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.188616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T16:59:50.615875Z digest=sha256:13c0ceed0eba3941289b6a9acc7ee0a924704f9f2ec21f338215b77f0778a982

Observation 470b5899-a398-4a91-b2f1-467d9bdc876a · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T04:34:50.786808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:34:50.786808Z digest=sha256:900892e17fcb3427c83cec39991596c36c6061b1919db5c24a0a60d7bc7220cb