Pith. sign in

Paper Citation Record · LEDGER

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

As of 13 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2608.08067.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08067 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:33:55.288126Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfec3e00-e796-406a-a40e-058c645abc9b · outbound

This paper cites arXiv preprint arXiv:2509.12508 (2025).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv preprint arXiv:2509.12508 (2025)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.203485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.203485Z digest=sha256:7c0cb42e783358d9fbf76e393f2f7fefd143f64f0c63ea826df298ac1bf9675d

Observation af35aa75-a3a9-409b-b5a1-59d5a5b76262 · outbound

This paper cites arXiv preprint arXiv:2509.22727 (2025).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv preprint arXiv:2509.22727 (2025)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.207732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.207732Z digest=sha256:86466196eef1a17a58e16a9f99036d2c54f526d40be137423f75b50f4597c687

Observation 97f4150d-0dd0-44b8-bb7b-047440424528 · outbound

This paper cites MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.211399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.211399Z digest=sha256:0eaacb98f906c6677c1ae676ed2267097b952f39d7a99e8abb3c1250504b7d59

Observation 200fa976-1e40-486f-a181-b679605b507b · outbound

This paper cites In: ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:33:55.215488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.215488Z digest=sha256:3126af46c8447146019f419d735a1460a18befcd3bdca2c6496b95b2a72052ad

Observation d4013e7a-d846-4835-ba93-cafe56731c97 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.219847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.219847Z digest=sha256:259899b15da01895d6fb18ecf08f2ec25fd685d58ee004a22af8cce397542c72

Observation 4ea41276-3ab0-4160-b57d-4a5f95c50811 · outbound

This paper cites Kimi-Audio Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Kimi-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.224384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.224384Z digest=sha256:d3223016e1e1665b17660d0e0a810c6659e4dee32bb1ddfd2dd2e31ba1841ec0

Observation 60815fdb-f522-4f01-84dc-8d8f7a04abb2 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.228960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.228960Z digest=sha256:7f32c8f75dabc27afcd1b0c54eed176c30db5edd6a2480f10df2b355870fe925

Observation d4c29cde-fb73-444f-aaa9-d1428a2a8819 · outbound

This paper cites In: International Conference on Learning Representations.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: International Conference on Learning Representations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:33:55.866459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:33:55.232967Z digest=sha256:4b13eaa8ba390f819313ece94778c097ed7365784b6d4493a838119d8a908b9b

Observation 1a96edab-e63f-4207-9307-e4acc974e0e7 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence40, 31627– 31635 (Mar 2026).https://doi.org/10.1609/aaai.v40i37.40429.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Proceedings of the AAAI Conference on Artificial Intelligence40, 31627– 31635 (Mar 2026).https://doi.org/10.1609/aaai.v40i37.40429

Reference 9

Resolution
verified exact
doi, observed 2026-08-12T00:33:55.339379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:33:55.235952Z digest=sha256:7a63657f49c3626548c1a896edbf9933de7b22dcb759621d06333de7a341c229

Observation a4b2af8b-8382-42fe-b922-fc84c899a034 · outbound

This paper cites Decoupled Weight Decay Regularization.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Decoupled Weight Decay Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.239061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.239061Z digest=sha256:0116d6c1fc4bc02b03ff704c4547ff76824e64350b6ce5b0198d8d5508d942e1

Observation 4d6f9104-32ed-4db6-8084-74f12415beea · outbound

This paper cites In: SC20: international conference for high performance computing, networking, storage and analysis.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: SC20: international conference for high performance computing, networking, storage and analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:33:55.853967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:33:55.242230Z digest=sha256:0a7d9bd71a247206238e29d1300d97c8d9f4eada8092e3f449989ce7524dec17

Observation efa11eb2-c6d0-42ce-855f-7586535084cf · outbound

This paper cites In: Proceed- ings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: Proceed- ings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.245198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.245198Z digest=sha256:cdd40cd0cdd0e27e58b8a900f88716fab9c3db11d6d48ae3e7187b511fe0840a

Observation 6bcb7b3d-2913-490a-b2e7-5baeaf936945 · outbound

This paper cites an unresolved cited work.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.248308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.248308Z digest=sha256:398e69084aaa8d1b447fc7522533775b22285e91ac36f64397f88e18d716edc1

Observation 97497c93-ea4a-4ecd-b488-ea0dfa8c8a84 · outbound

This paper cites Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.251451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.251451Z digest=sha256:92e0ffc2cbaa4e090b9c2f33d6d4f665899222723d478c497287ff316a819e5f

Observation e7814735-ac0f-48ae-8d53-84db3da3f315 · outbound

This paper cites Shu et al.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Shu et al

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:33:55.834988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:33:55.254938Z digest=sha256:1dfef06ac199355a4f4c08fe24a04b6262045625c380f518fe94270cf445f2a5

Observation cf630a7e-0095-48bc-8e8f-3b22c886bd23 · outbound

This paper cites arXiv preprint arXiv:2511.15848 (2025).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv preprint arXiv:2511.15848 (2025)

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.258278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.258278Z digest=sha256:633557fc2e272288530d44646a4c476af78afc56b087908adb463d56e258df49

Observation 8cb1243f-a25c-4aee-a581-73a6b85021d9 · outbound

This paper cites In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:33:55.824173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:33:55.261622Z digest=sha256:a261f5177fdcd892b960dc39af1de343311b550e0fda504e9029fb68bc358825

Observation 1eeb2fa5-51fe-4bc6-848d-7cab782ed022 · outbound

This paper cites Step-Audio 2 Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Step-Audio 2 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.264846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.264846Z digest=sha256:a26c7b7ab5db275c992536dc028b77a128909896a466b4ce70493c4a1d9240bd

Observation 8f00c36c-71ac-40ff-8c5e-6f0b1e3bedcc · outbound

This paper cites Qwen2.5-Omni Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Qwen2.5-Omni Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.268346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.268346Z digest=sha256:253b219e34406680128f338d95f8bf3cc341c37e22d2a4d458a83ba8bdb4a2e5

Observation 06de46fe-f455-487e-98d0-2453476583d0 · outbound

This paper cites Qwen3-Omni Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Qwen3-Omni Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.271928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.271928Z digest=sha256:6119d32b76eeb8d5d2b341b74510f07f20bfdc0c58b29b363a46ffde6e794bc3

Observation be9c1c1d-ba22-469e-8519-d2aa8261732f · outbound

This paper cites arXiv preprint arXiv:2603.10420 (2026).

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv preprint arXiv:2603.10420 (2026)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.276428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.276428Z digest=sha256:a0b95a8efce4e2ed9ec3917c0974f46448d3f51fd61b62b47853dfe6c48e5809

Observation ecb014c2-2379-4070-8554-663f4e7325d7 · outbound

This paper cites Qwen3 Technical Report.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Qwen3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.279827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.279827Z digest=sha256:1d849f00c44cfdd13685d34e56c85a1f50504aaa2688ca44d11cf9979820c213

Observation c648a9b2-30be-496d-9dd7-ddc0d543d6e7 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.283870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.283870Z digest=sha256:ffdbfc29f0a42aa68beb62f0265f202bd99a8b366f9cce8a6148866a781631cf

Observation fc5570f9-6f3e-4de1-a5c6-7576b8b645f1 · outbound

This paper cites EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:33:55.353666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:33:55.288126Z digest=sha256:9dc0cfb4d21332c9815b1260137d07330bf4512b3bfe722cfea34d3abc217914

Pith citing papers

No inbound Pith citation observations are available.