Pith. sign in

Paper Citation Record · LEDGER

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

As of 13 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2608.08067.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08067 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:33:55.288126Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfec3e00-e796-406a-a40e-058c645abc9b · outbound

This paper cites arXiv preprint arXiv:2509.12508 (2025).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv preprint arXiv:2509.12508 (2025)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.203485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.203485Z digest=sha256:58cd2526053b1f1e57e416ea2980148ab23dd7909edef959622311b00076788d

Observation af35aa75-a3a9-409b-b5a1-59d5a5b76262 · outbound

This paper cites arXiv preprint arXiv:2509.22727 (2025).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv preprint arXiv:2509.22727 (2025)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.207732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.207732Z digest=sha256:5eb84e826f9c825ede3345b20edbe7dd681fda08c8dad63a5f78a4bf1096c4c5

Observation 97f4150d-0dd0-44b8-bb7b-047440424528 · outbound

This paper cites MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.211399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.211399Z digest=sha256:f8ac28158ff120cca9230f89a6b293e4e801ce47abf4c6707f1a76a9530df709

Observation 200fa976-1e40-486f-a181-b679605b507b · outbound

This paper cites In: ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:33:55.215488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.215488Z digest=sha256:080033fe8ec3d85e8ae4638a8f594748eac11a6d81b17fb30f0ac7e217b2fd64

Observation d4013e7a-d846-4835-ba93-cafe56731c97 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.219847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.219847Z digest=sha256:a405ea7d29ddc6dcca8c6114e485b8d59e1f306eb55c65568709f39b93631ca9

Observation 4ea41276-3ab0-4160-b57d-4a5f95c50811 · outbound

This paper cites Kimi-Audio Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Kimi-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.224384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.224384Z digest=sha256:0bc7d06a0a5005dc9fa5883f7d100e2f2612aaa033269fe544c407ff8df87fe4

Observation 60815fdb-f522-4f01-84dc-8d8f7a04abb2 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.228960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.228960Z digest=sha256:108cb878efb89a94592b99ed7e0ff7845e08575719e346778f2e877725d3a6bc

Observation d4c29cde-fb73-444f-aaa9-d1428a2a8819 · outbound

This paper cites In: International Conference on Learning Representations.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: International Conference on Learning Representations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:33:55.866459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:33:55.232967Z digest=sha256:a89b3794a41647424b92b4ec9b2f557e4c1aec280f883fdff1bd307de576f526

Observation 1a96edab-e63f-4207-9307-e4acc974e0e7 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence40, 31627– 31635 (Mar 2026).https://doi.org/10.1609/aaai.v40i37.40429.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Proceedings of the AAAI Conference on Artificial Intelligence40, 31627– 31635 (Mar 2026).https://doi.org/10.1609/aaai.v40i37.40429

Reference 9

Resolution
verified exact
doi, observed 2026-08-12T00:33:55.339379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:33:55.235952Z digest=sha256:a0a218dff1a16e70bd3b8016d15358e594f80aa29928e543828f1293156d59ac

Observation a4b2af8b-8382-42fe-b922-fc84c899a034 · outbound

This paper cites Decoupled Weight Decay Regularization.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Decoupled Weight Decay Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.239061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.239061Z digest=sha256:b736d05e6b7dbf277d0ea18aca392e9fa0a0e85967f0727a0eb9b8988cfdc545

Observation 4d6f9104-32ed-4db6-8084-74f12415beea · outbound

This paper cites In: SC20: international conference for high performance computing, networking, storage and analysis.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: SC20: international conference for high performance computing, networking, storage and analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:33:55.853967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:33:55.242230Z digest=sha256:494008b86aaea1ccc244c605291673f89ee0c764327745b71f94782191d9390d

Observation efa11eb2-c6d0-42ce-855f-7586535084cf · outbound

This paper cites In: Proceed- ings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: Proceed- ings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.245198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.245198Z digest=sha256:9e80c685d17b44f9196636ff284ee4e9d4f3a67fb0eb4bd5801ec2f1d43d234d

Observation 6bcb7b3d-2913-490a-b2e7-5baeaf936945 · outbound

This paper cites an unresolved cited work.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.248308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.248308Z digest=sha256:8b8bde4fb83c0629e9901210d6028c963174e72a3bbedd302dd45e87121983de

Observation 97497c93-ea4a-4ecd-b488-ea0dfa8c8a84 · outbound

This paper cites Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.251451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.251451Z digest=sha256:7c5ba7a7eeaa10ebf8ef3fe6b6bee63dd43a5ee47fe4dbc33a688abb6da9b131

Observation e7814735-ac0f-48ae-8d53-84db3da3f315 · outbound

This paper cites Shu et al.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Shu et al

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:33:55.834988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:33:55.254938Z digest=sha256:313d1b0e3caf129a24e13bde180ec994c2f1abf0207a68cdc435915954d9ddca

Observation cf630a7e-0095-48bc-8e8f-3b22c886bd23 · outbound

This paper cites arXiv preprint arXiv:2511.15848 (2025).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv preprint arXiv:2511.15848 (2025)

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.258278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.258278Z digest=sha256:52427ac72c9a1955f090f64f1c25e23c2a24584ed8fd939307c95db2c63a2936

Observation 8cb1243f-a25c-4aee-a581-73a6b85021d9 · outbound

This paper cites In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:33:55.824173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:33:55.261622Z digest=sha256:9a57bec42051454ad6732438abf81b7d0fdef58ae830e863d048bd79d5e97ba8

Observation 1eeb2fa5-51fe-4bc6-848d-7cab782ed022 · outbound

This paper cites Step-Audio 2 Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Step-Audio 2 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.264846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.264846Z digest=sha256:a6697ef4e567739be0270c8c44cd8cc36699aaa8d37d8ec71fe47f59739975dd

Observation 8f00c36c-71ac-40ff-8c5e-6f0b1e3bedcc · outbound

This paper cites Qwen2.5-Omni Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Qwen2.5-Omni Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.268346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.268346Z digest=sha256:d82fad5a92a94bfd1d62f2a82fcac5e8ebb9c5bdc1b1b43a3cf29b29439ba670

Observation 06de46fe-f455-487e-98d0-2453476583d0 · outbound

This paper cites Qwen3-Omni Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Qwen3-Omni Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.271928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.271928Z digest=sha256:e896b57b88c2c6ae1a63ec8a1cdc4da95e2565c43090873a40002634b32786e9

Observation be9c1c1d-ba22-469e-8519-d2aa8261732f · outbound

This paper cites arXiv preprint arXiv:2603.10420 (2026).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv preprint arXiv:2603.10420 (2026)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.276428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.276428Z digest=sha256:5545ce4ca546612b97ead9e232cead36af49b64a56fae0eee2610c370796427b

Observation ecb014c2-2379-4070-8554-663f4e7325d7 · outbound

This paper cites Qwen3 Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Qwen3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.279827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.279827Z digest=sha256:b3c1601f933dfffe2e904184f87f1766dc240c5c1f4f9b8dfef4a9d6f5c0b561

Observation c648a9b2-30be-496d-9dd7-ddc0d543d6e7 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.283870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.283870Z digest=sha256:af41eccbdd1aaf9802af3d74d12b30defc4c7064db8d158d0d4a10054c444136

Observation fc5570f9-6f3e-4de1-a5c6-7576b8b645f1 · outbound

This paper cites EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:33:55.353666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:33:55.288126Z digest=sha256:b6d3e9ef349e907dea97c0d3d6b11c89f88d6fb506456a07da1b0ebbca6687a3

Pith citing papers

No inbound Pith citation observations are available.