Pith. sign in

Paper Citation Record · LEDGER

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

As of 6 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2507.18119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18119 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:41:57.488433Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:03:09.632009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T20:59:47.893961Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb64e32c-ff90-48d9-b9f2-575e4fb635e2 · outbound

This paper cites On the landscape of spoken language models: A comprehensive survey,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness On the landscape of spoken language models: A comprehensive survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.914459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.365862Z digest=sha256:6b9ae6752361ae3b579763e2a0566f072e17ef5e9fbd231f288aee0d980cdf24

Observation 68bbeba8-9c7c-467d-b81d-bd99ce2c822f · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.901450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.371356Z digest=sha256:75805575e1031aacc088a2809a8bf0741948997a750bbe53a992112c523db51e

Observation 24041366-a50b-40ad-a380-5a6e1234c62c · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Moshi: a speech-text foundation model for real-time dialogue,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.888581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.376159Z digest=sha256:cd75e67b4635810385ce9fe9ae476309a8b57dd958d304764695a98e05349bc6

Observation 72c5bfa6-bcc3-430b-a2bc-3d0e4a717200 · outbound

This paper cites Speechgpt 2.0-preview,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Speechgpt 2.0-preview,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.875689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.381247Z digest=sha256:29f2c00bee77583edbcc54bfff8ba1c1b31a7e6f6a07b8089fe7fa68c4cfefc2

Observation 938b6b44-702d-478d-a7f9-d18e29f44c28 · outbound

This paper cites Llama-omni: Seamless speech interaction with large language models,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Llama-omni: Seamless speech interaction with large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.862931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.385844Z digest=sha256:0874abba1e49113a275a49612017d38080e07c01a6a77da62c27fb3092cd27a2

Observation e72106f9-554c-4ef0-8642-7a89c502a69f · outbound

This paper cites Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.849408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.390739Z digest=sha256:8997a5bb8b01276ff06d42d1c174f10f031b1842eaec1e52d4d8499b8827d265

Observation 36bed796-faa9-4a2e-96fc-0c967461f38a · outbound

This paper cites Slam-omni: Timbre-controllable voice interaction system with single-stage training,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Slam-omni: Timbre-controllable voice interaction system with single-stage training,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.836513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.396459Z digest=sha256:47394a061f9f77b343208f2131fe6c70c8cb34b2b7d69f0c663ef6916c1a1cb0

Observation 02c73c13-ce25-43d2-8f96-5a2a7c3f9521 · outbound

This paper cites Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.823488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.400457Z digest=sha256:87119239929bb90eff86985a770c1aa31eca899182a89c278e3e38ffaf3cf90a

Observation cc859baf-0b64-426d-a228-f8fba81a3694 · outbound

This paper cites Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live streaming on your phone,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live streaming on your phone,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.810776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.404284Z digest=sha256:15939020c4588d7a566cdcdb6042fef25d9a3f78333388c81c2cc6703c4a670f

Observation cdb7173f-a781-48aa-8211-75e898235a4a · outbound

This paper cites Baichuan-omni-1.5 technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Baichuan-omni-1.5 technical report,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.796923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.408621Z digest=sha256:57399894a96b82678af91bb8fb2ea02bf25320d707f599ae50a5bc74bd286286

Observation 000ff57d-2b87-4200-8928-eda716400dae · outbound

This paper cites Salmonn-omni: A codec-free LLM for full-duplex speech understanding and generation,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Salmonn-omni: A codec-free LLM for full-duplex speech understanding and generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.782628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.412961Z digest=sha256:f29c7c337f48904dc52368a6012733ccf94d2b83dfcc594112193e99764ab4c6

Observation d6181bd6-0b70-41ca-a7aa-c952f014718c · outbound

This paper cites Minmo: A multimodal large language model for seamless voice interaction,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Minmo: A multimodal large language model for seamless voice interaction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.768963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.417459Z digest=sha256:c093267a4afab62ea3956771346b379aa920ad1c69d47195b990c4749a744b1f

Observation 9e6aff3e-aff5-4316-b407-9cb0517125a5 · outbound

This paper cites Qwen2.5-omni technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Qwen2.5-omni technical report,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.755597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.421720Z digest=sha256:48c49d70f8f481b8cbc3813aabf0f725d8fb20dcc5d5e9b82e9d486d07781e37

Observation b5cd58e1-1896-4480-91a6-0998d2991aa4 · outbound

This paper cites Step-audio: Unified understanding and generation in intelligent speech interaction,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Step-audio: Unified understanding and generation in intelligent speech interaction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.742003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.426101Z digest=sha256:6c7753c4c49873d2f3596e8820f2bd4649d19cc52ab3e36e095b168261a61ef9

Observation b5606be8-d984-47f7-be92-d774c481c6c4 · outbound

This paper cites Step-Audio-AQAA: a fully end-to-end expressive large audio language model,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Step-Audio-AQAA: a fully end-to-end expressive large audio language model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.727510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.430435Z digest=sha256:7c26ea0e74a4312f3c67dad86e0ba65fc453657d8e385dc3e8879ea879396192

Observation 7c8a943b-d78f-4531-92a9-a43779cb87a7 · outbound

This paper cites Llama-omni2: Llm-based real-time spoken chatbot with autoregressive streaming speech synthesis,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Llama-omni2: Llm-based real-time spoken chatbot with autoregressive streaming speech synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.713857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.435030Z digest=sha256:cf2953e11e8653a664bbcb3ae1d5d9803ccb408414f0cbf034700f3857f1cdf3

Observation f7d4abe7-6ffa-4d74-8890-cef3efe68d10 · outbound

This paper cites Deeptalk: Towards seamless and smart speech interaction with adaptive modality-specific moe,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Deeptalk: Towards seamless and smart speech interaction with adaptive modality-specific moe,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.698773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.439127Z digest=sha256:e45e19689441ebcf0075451c96bfff452ac94c99a5c6bf65ea00fcceff85843f

Observation 16f671c4-e59e-470c-aac9-dc8dc68bbab7 · outbound

This paper cites BoSS: Beyond-semantic speech,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness BoSS: Beyond-semantic speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.685148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.443222Z digest=sha256:5ce48ad28d01539e4ce406acea0a3d406f91e7087dd75e2155ec16e6218b6f92

Observation 29a2c081-b8f8-430c-a9cf-a6f7328a828f · outbound

This paper cites V oila: V oice-language foundation models for real-time autonomous interaction and voice roleplay,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness V oila: V oice-language foundation models for real-time autonomous interaction and voice roleplay,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.670291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.447315Z digest=sha256:3e7b74cd5be3339ea644efac878c35ca60b717d0276d2dc3180f5fa6905a36ce

Observation 9f8af00c-cc09-4c45-bfea-071d41d386b0 · outbound

This paper cites Kimi-audio technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Kimi-audio technical report,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.657231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.451156Z digest=sha256:37a304a180486268d3dc025f78d0105101daa8b49aee567165a2538b68736535

Observation 0243e4ed-6eda-4fc8-b4ab-4ada7e33f86d · outbound

This paper cites GOAT- TTS: llm-based text-to-speech generation optimized via A dual-branch architecture,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness GOAT- TTS: llm-based text-to-speech generation optimized via A dual-branch architecture,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.643553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.455369Z digest=sha256:13dc7383c61c7913ff53df26be4d41dfda9be4018a60d1e17a82706f9c0c80d1

Observation 1bb63ac0-1f4d-412c-a07b-5700418282e0 · outbound

This paper cites TELEV AL: A dynamic benchmark designed for spoken language models in chinese interactive scenarios,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness TELEV AL: A dynamic benchmark designed for spoken language models in chinese interactive scenarios,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.630056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.459831Z digest=sha256:28b27bad75e88676403e9742a3459ba0f10ba448f9bd48725436f39dc9cec40a

Observation 4b62431f-b158-4d8f-b63a-611371451226 · outbound

This paper cites Qwen2-audio technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Qwen2-audio technical report,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.616374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.464545Z digest=sha256:80c751dd3e03cb22776edf99603f5595219a089957e88208b9a0b2f6e1ef7695

Observation 4385ee2b-2e76-42bf-8e1e-6c072eb367d2 · outbound

This paper cites Baichuan-Audio: A unified framework for end-to-end speech interaction,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Baichuan-Audio: A unified framework for end-to-end speech interaction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.601830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.468284Z digest=sha256:7b68added753c520b24d2e2a1ac4419369a98dd4a244436e6a47c2fa2a8cfd44

Observation b846dc80-2c58-42c1-adda-536d03c04710 · outbound

This paper cites Telechat technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Telechat technical report,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.586620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.472468Z digest=sha256:30e169bd9ddd496a61f85544ae781c75a6e0c664947abcd113955c4fe3af8947

Observation a872b40c-1ecb-4efb-8603-fad8c29bbf85 · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Audiochatllama: Towards general-purpose speech abilities for llms,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.572007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.476480Z digest=sha256:306acc31a02d4de75d449bd2e5b7c7073fdc2377a09fed5d028dfd337959547f

Observation f5806a61-abdf-41cb-b52c-e4bb51384651 · outbound

This paper cites BLSP: bootstrapping language-speech pre-training via behavior alignment of continuation writing,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness BLSP: bootstrapping language-speech pre-training via behavior alignment of continuation writing,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.557553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.480339Z digest=sha256:299bf9b52d365c1d9edb0890fe80fa0d359e115565391b81cc8075de977f5d10

Observation 6db03b0e-9444-4583-a2ec-65646a9a624d · outbound

This paper cites Wav2prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based llms,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Wav2prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based llms,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.542387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.484253Z digest=sha256:7c303974c0dddfaada7ac5453beaf2959f03bbe140015506ce55d7557d6d5923

Observation 3bac4abc-912c-46b0-87bb-31a4d20a8f0f · outbound

This paper cites DeSTA2: Developing instruction-following speech language model without speech instruction-tuning data,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness DeSTA2: Developing instruction-following speech language model without speech instruction-tuning data,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.527293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:41:57.488433Z digest=sha256:324ff64feedb3375021c71aaa30232338414c3da86cb2ea0e07e0073bead033e

Pith citing papers

Observation 114b1b5d-ca08-4e4f-99be-214cde165239 · inbound

BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation cites this paper.

BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

Reference 2008

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:59:47.953809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T20:59:43.785209Z digest=sha256:4a0e22795844a55c59bcfd333b85ee370e3b5c2ead563680a59406ac0ae82fec

Observation 5affa23d-c094-4c16-896c-d006ce8796cc · inbound

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue cites this paper.

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-05T21:03:09.632009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:03:09.632009Z digest=sha256:0d1430635c7e4f07ed7fec1a262afd3231986299fc36458cd0dd7ac675332d8d