Pith. sign in

Paper Citation Record · LEDGER

Voice Memory for Agentic Speech Recognition

As of 6 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2607.26410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26410 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:37:24.396537Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aaf56af3-0908-473b-b7f0-d6bb8ee556cd · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

Voice Memory for Agentic Speech Recognition GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.132795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.132795Z digest=sha256:42cc101ad1bad2bf2f75347dedfe9baf3e135f572aad86d62e5abfe7652235e1

Observation 8ccd3b20-999a-4b55-9786-cab5e26c4be4 · outbound

This paper cites Bringing contextual information to google speech recognition.

Voice Memory for Agentic Speech Recognition Bringing contextual information to google speech recognition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.138996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.138996Z digest=sha256:a5207dadf51dac6e2c66eb4c8cd98210879f81654ea43e7a3313bcbaa7623387

Observation 3461982e-59d1-4325-ab0b-8ca4d5d00424 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Voice Memory for Agentic Speech Recognition wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.143907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.143907Z digest=sha256:0ae258daeeefa1572a3d9592d52db1c143f0413b6e864a4ffdf027943ddacf69

Observation d694f629-a56a-4990-86ba-a501dc257181 · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al.

Voice Memory for Agentic Speech Recognition Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.148371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.148371Z digest=sha256:f3dc8860f846ff56807ad28726cb4b9d2279645914389978b4a679d62f94d99c

Observation 07f83cb1-9511-46ff-b4aa-655e95e629f8 · outbound

This paper cites Text-only adaptation in llm-based asr through text denoising.arXiv preprint arXiv:2601.20900, 2026.

Voice Memory for Agentic Speech Recognition Text-only adaptation in llm-based asr through text denoising.arXiv preprint arXiv:2601.20900, 2026

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.153284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.153284Z digest=sha256:2174d6e1ea7aad07f9fbbf8ebf9d01051f1d7ff9764b50aa4185cd921ae71cc4

Observation 065a9479-034a-445a-a55d-a38bcd3b7ec6 · outbound

This paper cites HyPoradise: An open baseline for generative speech recognition with large language models.

Voice Memory for Agentic Speech Recognition HyPoradise: An open baseline for generative speech recognition with large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.157509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.157509Z digest=sha256:a203b431552b78a04843e10ef68f727b53cfe92379b68f4e485be0a262c854d0

Observation 27475754-cec9-40d5-996f-e4ef759b8482 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Voice Memory for Agentic Speech Recognition Moshi: a speech-text foundation model for real-time dialogue

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.162267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.162267Z digest=sha256:64c4e4713021fb505f9baf3b64be93ccfdb5a8900d01f6496fc4fde6dc8e1a57

Observation 20912ae0-15f5-4a0d-846f-2cfeee225fed · outbound

This paper cites Lipger: Visually-conditioned generative error correction for robust automatic speech recognition.

Voice Memory for Agentic Speech Recognition Lipger: Visually-conditioned generative error correction for robust automatic speech recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.166387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.166387Z digest=sha256:29d2b55a72b3c4c88b786184f8021c5614c4408e81a0a4a5f0df5c27156c9146

Observation b78b7b3f-a79f-4e78-b60e-5dfa5910f0cc · outbound

This paper cites Words and voices: episodic traces in spoken word identification and recognition memory.Journal of experimental psychology: Learning, memory, and cognition, 22(5):1166, 1996.

Voice Memory for Agentic Speech Recognition Words and voices: episodic traces in spoken word identification and recognition memory.Journal of experimental psychology: Learning, memory, and cognition, 22(5):1166, 1996

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.170554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.170554Z digest=sha256:2f6e0d3396b9ee05508d3398cac0990ee52f5ecd54996d3d8b4004058b2885c5

Observation 59d3e9b6-1e04-437e-8999-b489dd4b93af · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition.

Voice Memory for Agentic Speech Recognition Conformer: Convolution-augmented transformer for speech recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.174413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.174413Z digest=sha256:e18e24c6c230b2c1f883a356b95712996b376c8302b88cb35938cc0bae2269d0

Observation 74716500-4813-4c5f-838b-bef3fe97579d · outbound

This paper cites The sapir-whorf hypothesis.Language in culture, 4:92–105, 1954.

Voice Memory for Agentic Speech Recognition The sapir-whorf hypothesis.Language in culture, 4:92–105, 1954

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.178316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.178316Z digest=sha256:2eeb45272b8d4f1dfc8f097626665da8edf52204dd60325bd87c3c9c540a02da

Observation 8dda4e6a-72d0-4a4f-8498-25ef102bd7a5 · outbound

This paper cites Let’s fuse step by step: A generative fusion decoding algorithm with llms for robust and instruction-aware asr and ocr.

Voice Memory for Agentic Speech Recognition Let’s fuse step by step: A generative fusion decoding algorithm with llms for robust and instruction-aware asr and ocr

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.182179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.182179Z digest=sha256:46c72749e8f94c9131c96d01cd42109515c3a8da565aa5aaa6e537610b6a1fe5

Observation 4f7d2f51-dac6-49eb-b430-a80a0f6d00c4 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.

Voice Memory for Agentic Speech Recognition HuBERT: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.186304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.186304Z digest=sha256:b6b3845de84b6bb8caccfa5001895deb0be89a115f6e62c4b25201f111c81262

Observation 3a23f3fe-b598-48f1-bf9b-e232856ab499 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Voice Memory for Agentic Speech Recognition Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.190225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.190225Z digest=sha256:21222fe6d68f91de30cf36daccce490d63c722f144d505a9c2bfd485051831bb

Observation 0e37dfb7-6e7d-4fd6-bdc9-5e934c9f1c60 · outbound

This paper cites Scaling up deliberation for multilingual asr.

Voice Memory for Agentic Speech Recognition Scaling up deliberation for multilingual asr

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.194118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.194118Z digest=sha256:e3d78e82f3b96c40e0c78020ffb33cacdf24cf8f42abf9d1c199446377fdf43f

Observation a2ad02d0-90cf-4c3d-a9dd-500ad38aa435 · outbound

This paper cites Chain-of-thought prompting for speech translation.

Voice Memory for Agentic Speech Recognition Chain-of-thought prompting for speech translation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.197901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.197901Z digest=sha256:e2fde3daa0e4515a981e455d25d48a634ab3c910d2607ba0bd8f276420c5bd40

Observation 6891f0df-8a3b-4213-98ee-8370ff360f2c · outbound

This paper cites Subjective comparison and evaluation of speech enhancement algorithms.

Voice Memory for Agentic Speech Recognition Subjective comparison and evaluation of speech enhancement algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.202529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.202529Z digest=sha256:3ee6c64f8b18bc2ede8889afec5719188ef1b1248573e441330b17f2a34a0c3d

Observation bfa8f024-9373-4744-ae7b-124af05e5ed4 · outbound

This paper cites Large language models are efficient learners of noise-robust speech recognition.

Voice Memory for Agentic Speech Recognition Large language models are efficient learners of noise-robust speech recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.206425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.206425Z digest=sha256:f80369ccd861049b0c970c93a9d0902bc1f57f0971df2c4569dc9b6a7de64a86

Observation dc3a39b5-7cb1-4473-bfb2-7a03b2de8fce · outbound

This paper cites GenTranslate: Large language models are generative multilingual speech and machine translators.

Voice Memory for Agentic Speech Recognition GenTranslate: Large language models are generative multilingual speech and machine translators

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.210422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.210422Z digest=sha256:1c26346eec119fcde94eb1fc998825d8d41a6b2963fe8f2cd112abe746f2b533

Observation 89b5d4f3-5ebc-40ac-b54e-9fcb33ce1ef0 · outbound

This paper cites DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction.

Voice Memory for Agentic Speech Recognition DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.214630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.214630Z digest=sha256:4859393086963d00a0bfe4f3b57612efffa58d2db7cad78224fcbfacacdc9e0c

Observation 00af5936-5fdb-4d80-9076-05c9b61f9c98 · outbound

This paper cites Evaluating open-source asr systems: Performance across diverse audio conditions and error correction methods.

Voice Memory for Agentic Speech Recognition Evaluating open-source asr systems: Performance across diverse audio conditions and error correction methods

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.219368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.219368Z digest=sha256:b33fb2ab93098c3ed49391556c173568c57944a5b50b86b0bc5fc9558b3bc4bc

Observation 3867057d-38b2-445a-87ae-604f78390873 · outbound

This paper cites Contextual RNN-T for open domain ASR.

Voice Memory for Agentic Speech Recognition Contextual RNN-T for open domain ASR

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.223494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.223494Z digest=sha256:416e568dd8653dbc566688118c9844a606f34d4e285a19d4bb25cfed0ca0733d

Observation b4a305c3-4c14-4fde-95a3-359500ca9ae0 · outbound

This paper cites Language, Speech and Communication.

Voice Memory for Agentic Speech Recognition Language, Speech and Communication

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.227749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.227749Z digest=sha256:2c6b73f414526521adaafdd2de2b1e4b04765344ebeac1aa4aa361e97f819126

Observation 2f933df7-0693-4681-81df-88eee78d2331 · outbound

This paper cites Two heads are better than one: Audio-visual speech error correction with dual hypotheses.

Voice Memory for Agentic Speech Recognition Two heads are better than one: Audio-visual speech error correction with dual hypotheses

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.231756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.231756Z digest=sha256:9f69a0c6beca38de1977d90a0638c72fe6bfe45e3585c7f601329d3de576c7b0

Observation 0c20d1e4-5b1b-4885-844f-537575bb3d94 · outbound

This paper cites Semantic distance: A new metric for ASR performance analysis towards spoken language understanding.

Voice Memory for Agentic Speech Recognition Semantic distance: A new metric for ASR performance analysis towards spoken language understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.235664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.235664Z digest=sha256:3a12f1f8e98eaae0dba5dd343949f7977a2e3d5ee9e02878570a885d08f9cf42

Observation e52b8d79-183e-4cf1-894a-2fcd6af005f9 · outbound

This paper cites Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs.

Voice Memory for Agentic Speech Recognition Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.239501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.239501Z digest=sha256:9fab040000e07d7ca762e576bdaa712c4f86478b77850f1329f19c067893e7ea

Observation 113fa126-2834-4a4f-8613-0649d816cd4b · outbound

This paper cites Interaction models: A scalable approach to human-ai collabora- tion.Thinking Machines Lab: Connectionism, May 2026.

Voice Memory for Agentic Speech Recognition Interaction models: A scalable approach to human-ai collabora- tion.Thinking Machines Lab: Connectionism, May 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.243702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.243702Z digest=sha256:79109bc6a70faeb863aaface788924127e57f44155928bba46795cc7b902a4ae

Observation 22d06815-a0ce-4217-8d19-4078087f5db7 · outbound

This paper cites MiniMax Sparse Attention.

Voice Memory for Agentic Speech Recognition MiniMax Sparse Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.248029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.248029Z digest=sha256:338dcfc4e633961b9a17a84ae3ddad28ef4602a1aae77419b95144e2e8218537

Observation bef6a746-b4a6-472b-8ce5-5bbfcf9c7256 · outbound

This paper cites an unresolved cited work.

Voice Memory for Agentic Speech Recognition Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.252759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.252759Z digest=sha256:9c93531fe4101e20a97f9fca3a549bc96f52aef007a653649f86bc66c608d83f

Observation 24b3fdb6-fae2-44a9-8843-ab912977895f · outbound

This paper cites FastCorrect: Fast error correction with edit alignment for automatic speech recognition.

Voice Memory for Agentic Speech Recognition FastCorrect: Fast error correction with edit alignment for automatic speech recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.256835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.256835Z digest=sha256:eb8c9749630b83f9962b38452f55fb5ba9a41fff7bb3a5eaf922409894d3fdcd

Observation 4af70134-b8e9-4d5b-9e69-f541aa34e855 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Voice Memory for Agentic Speech Recognition The power of scale for parameter-efficient prompt tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.260749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.260749Z digest=sha256:64c38e57d962d3257ed28a44402bdbf42cd74b589ce2018ef6026e5cd0a8c356

Observation 96fa0818-26d1-4c76-9bf4-bf30f1185d31 · outbound

This paper cites Investigating asr error correction with large language model and multilingual 1-best hypotheses.

Voice Memory for Agentic Speech Recognition Investigating asr error correction with large language model and multilingual 1-best hypotheses

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.264731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.264731Z digest=sha256:6d03bc8d15a238957fefa97611b4e0cf88b70d316be2182732e062f7d1ca4ea2

Observation 583c4361-df41-4434-ae9e-38058de937bf · outbound

This paper cites Rethinking evaluation in ASR: Are our models robust enough? InProc.

Voice Memory for Agentic Speech Recognition Rethinking evaluation in ASR: Are our models robust enough? InProc

Reference 33

Resolution
verified exact
doi, observed 2026-08-01T16:39:07.709045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-01T16:37:24.268826Z digest=sha256:1517321045c7eebcd0eb8d562132c0c2779007fb89465072a840f5a0cf3b7cc7

Observation a279c5b1-6523-49d4-8190-9e393675b228 · outbound

This paper cites Neko: Cross-modality post-recognition error correction with tasks-guided mixture-of-experts language model.

Voice Memory for Agentic Speech Recognition Neko: Cross-modality post-recognition error correction with tasks-guided mixture-of-experts language model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.273275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.273275Z digest=sha256:2ca862cee92709661bee3bb7898b53db7032c16e74d47d753c1f0a082640b505

Observation 46b4055d-fd60-450f-b238-122cf2b88aee · outbound

This paper cites N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space.

Voice Memory for Agentic Speech Recognition N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.277296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.277296Z digest=sha256:bf1f633b62e75ca981db62b22f1c091e0358ff93f646beb945900d72a85f405f

Observation 83179833-5be6-4cd1-a3dc-c7817b560147 · outbound

This paper cites Typicality effects on memory for voice: Implications for earwitness testimony.Applied Cognitive Psychology, 25(1):29–34, 2011.

Voice Memory for Agentic Speech Recognition Typicality effects on memory for voice: Implications for earwitness testimony.Applied Cognitive Psychology, 25(1):29–34, 2011

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.282117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.282117Z digest=sha256:7f9124c971efd9937805a8ba89fc7109b159fa921f3525de3845bf4d4721303c

Observation e5524955-731e-4e8a-895a-6623f4e2ee95 · outbound

This paper cites Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills.

Voice Memory for Agentic Speech Recognition Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.286644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.286644Z digest=sha256:48d306625254e1be40aedacad4c8d062643b1a80fd7fc736d0f51f535d6267e0

Observation aee61dff-fc88-41e1-8859-652290aab18f · outbound

This paper cites Long-term memory for unfamiliar voices.The Journal of the Acoustical Society of America, 85(2):913–925, 1989.

Voice Memory for Agentic Speech Recognition Long-term memory for unfamiliar voices.The Journal of the Acoustical Society of America, 85(2):913–925, 1989

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.291191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.291191Z digest=sha256:c59b061b7c82c3b45a8166f2bc68127d7fe1f08e494120df0248b58115b83428

Observation 57f84aa3-5ca5-437a-a601-4ec402a47ac3 · outbound

This paper cites Sainath, Rohit Prabhavalkar, Anjuli Kannan, and Ding Zhao.

Voice Memory for Agentic Speech Recognition Sainath, Rohit Prabhavalkar, Anjuli Kannan, and Ding Zhao

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.295465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.295465Z digest=sha256:f857f3753bd05534fa4cd270d2ac758213754d8d1fbe8824436de8accefcc654

Observation ce5f10c9-3d89-450f-b309-6e579f78d280 · outbound

This paper cites Less is more: Accurate speech recognition & translation without web-scale data.

Voice Memory for Agentic Speech Recognition Less is more: Accurate speech recognition & translation without web-scale data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.299615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.299615Z digest=sha256:012662bcaaab383b088768e225471661ca11607d6b53eda3410856f2250dc069

Observation bd1a3348-bf53-4931-8fa7-1a6026897380 · outbound

This paper cites Qwen3 Technical Report.

Voice Memory for Agentic Speech Recognition Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.303661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.303661Z digest=sha256:178fe80984cdf1fc4a540199abf69bfc253ec96674c048d42f5dbe4356f10ea5

Observation fd4e9969-d628-4015-bbea-180467378d9d · outbound

This paper cites an unresolved cited work.

Voice Memory for Agentic Speech Recognition Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.307961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.307961Z digest=sha256:2c9090cf4ad9c94c21a7f5cada1604eb4fb559fd3302e2c65ea722f97aff3e75

Observation d7b1c692-eac7-434b-b97f-885a305d1de4 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Voice Memory for Agentic Speech Recognition Robust speech recognition via large-scale weak supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.312091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.312091Z digest=sha256:3b50ea3e7c8f1e6751659358738eb17b37113a9fca357f9de430339bccb2ab61

Observation 077198d2-65a2-46ea-b632-a0c96f460bfc · outbound

This paper cites Whispering LLaMA: A cross-modal generative error correction framework for speech recognition.

Voice Memory for Agentic Speech Recognition Whispering LLaMA: A cross-modal generative error correction framework for speech recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.316031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.316031Z digest=sha256:b450bd309f6b92d27a497d388941c35534aae2db699e4b07bdf091d938de19d6

Observation 2e4ad00a-23de-43a4-b356-90ba8bf1e061 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Voice Memory for Agentic Speech Recognition Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.320281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.320281Z digest=sha256:a2a5279a7c74c829b34368d1f03b4f2fd6534b926de284c0502ef5a832d55111

Observation 845d37c5-b8b0-484f-8cd5-c69b008851b3 · outbound

This paper cites Fast conformer with linearly scalable attention for efficient speech recognition.

Voice Memory for Agentic Speech Recognition Fast conformer with linearly scalable attention for efficient speech recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.324366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.324366Z digest=sha256:271603b763f1f427979b4ad7f92dbb2e2f23198193502e84b3eba64ad61f7f05

Observation 70f70d70-95b0-41ed-95f9-f9c897b46af8 · outbound

This paper cites Nguyen, and Katrin Kirchhoff.

Voice Memory for Agentic Speech Recognition Nguyen, and Katrin Kirchhoff

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.328822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.328822Z digest=sha256:456e353b4a638d168aa8e2b704ce424cbc35cedc670e7c496e578fa5d6bdeeed

Observation 71d7e1cb-fe47-40df-a5f0-2759e3b58596 · outbound

This paper cites The effects of voice and visible speaker change on memory for spoken words.Journal of Memory and Language, 34(5):665–685, 1995.

Voice Memory for Agentic Speech Recognition The effects of voice and visible speaker change on memory for spoken words.Journal of Memory and Language, 34(5):665–685, 1995

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.333588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.333588Z digest=sha256:785c0222f723f7e3d51bc3960a5a685cc3b11f4e34e1ef7417a3d0b189a2f9e0

Observation 70c9d561-2872-45a4-9eca-e29d952b0316 · outbound

This paper cites Dialog act modeling for automatic tagging and recognition of conversational speech.Computational Linguisitics, 26(3):339, 2000.

Voice Memory for Agentic Speech Recognition Dialog act modeling for automatic tagging and recognition of conversational speech.Computational Linguisitics, 26(3):339, 2000

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.337785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.337785Z digest=sha256:f223d0bc85109938f22d415730c6ebb84564581de61b4db5ddade43589a1928d

Observation 551bee93-2e5b-4c72-ade8-c932bc51bbe5 · outbound

This paper cites WER we are and WER we think we are.

Voice Memory for Agentic Speech Recognition WER we are and WER we think we are

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.342441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.342441Z digest=sha256:c160bf47cf4e2f53ed3efef92313afd54ce8eebc475525324133b70dbe1faf0c

Observation 2e78cfaa-e8ee-4a7a-81f6-cdb36ffd8e67 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models.

Voice Memory for Agentic Speech Recognition Salmonn: Towards generic hearing abilities for large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.346467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.346467Z digest=sha256:f58f73955642c497830908ce7e7ab3fc57ec94978a149706612524ec222eea38

Observation ed815807-3c00-4f7d-907f-57c99e76a297 · outbound

This paper cites Interaction models: A scalable approach to human-AI collaboration, May 2026.

Voice Memory for Agentic Speech Recognition Interaction models: A scalable approach to human-AI collaboration, May 2026

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.350593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.350593Z digest=sha256:3bbb657b2182e2934d08fd0e17350f129f260fdde37c278177097766a3b1ea30

Observation cf218351-4c69-4227-a54f-feefa21a5451 · outbound

This paper cites Leveraging llms for post-ocr correction of historical newspapers.

Voice Memory for Agentic Speech Recognition Leveraging llms for post-ocr correction of historical newspapers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.354772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.354772Z digest=sha256:3755c8f46af8970f0f1c26edeeeb872db43036dc7839c7f2c74727ef3ab948ca

Observation 12a5a23d-35a5-45f7-b2e4-5991921249b6 · outbound

This paper cites Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception.

Voice Memory for Agentic Speech Recognition Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.358711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.358711Z digest=sha256:14eb345fdfabc97c49e2c5a63b99a1b0c0539f96c68750be9cea924954be9851

Observation b931952c-42e1-4b70-a83b-9a03bd5b5a6c · outbound

This paper cites Audio-Mind: An Auditable Agentic Framework for Audio Understanding.

Voice Memory for Agentic Speech Recognition Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.362798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.362798Z digest=sha256:d3f9fe45fcf389e21b539c3e56ade6c6c1da797a13a8d1a156a838fac9c06e32

Observation c3cd97e2-d91c-4957-bb7a-71159b7c7b44 · outbound

This paper cites SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models.

Voice Memory for Agentic Speech Recognition SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.367012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.367012Z digest=sha256:225225239f3a17ec25fb62ca0d5228e2252ffcd462c7dcb4b532d53f50bf7764

Observation 024b5c21-90a2-48ce-b60b-7b75ec15f51d · outbound

This paper cites Generative speech recognition error correction with large language models and task-activating prompting.

Voice Memory for Agentic Speech Recognition Generative speech recognition error correction with large language models and task-activating prompting

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.371144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.371144Z digest=sha256:77f5789e53f4b3b5b26a49baa96a2c085c4703b73e6f5798ff0b01dc96af29e4

Observation ff0518b6-a4bf-4af2-8623-41b031b56655 · outbound

This paper cites Large language models as optimizers.

Voice Memory for Agentic Speech Recognition Large language models as optimizers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.375189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.375189Z digest=sha256:1b4b414c0b8975367b72c084b1532669e889d1c268c3aee290fb1578a7abf424

Observation 7610b86e-b951-443d-b697-4723d06c2c3d · outbound

This paper cites SkillOpt: Executive Strategy for Self-Evolving Agent Skills.

Voice Memory for Agentic Speech Recognition SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.379151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.379151Z digest=sha256:ab91e5d9d4bc77d6481b463a8cf59e48be32b1f1cd5f4b045eca3289cb90397e

Observation b784d6f9-8749-48ba-9644-8727eba72c83 · outbound

This paper cites Covoger: A multilingual multitask benchmark for speech-to-text generative error correction with large language models.

Voice Memory for Agentic Speech Recognition Covoger: A multilingual multitask benchmark for speech-to-text generative error correction with large language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.383680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.383680Z digest=sha256:2449578286e5b2786f8296f410546fcce4112d77daa8582b0c5cf099993e9eb0

Observation 1864157d-29fb-41c2-80fb-d38345acbfca · outbound

This paper cites Mms-llama: Efficient llm-based audio-visual speech recognition with minimal multimodal speech tokens.

Voice Memory for Agentic Speech Recognition Mms-llama: Efficient llm-based audio-visual speech recognition with minimal multimodal speech tokens

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.388102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.388102Z digest=sha256:add608331bd30beffe0bb07666862a522c50bc3dd00ba1ebc956e280b8d1ef2a

Observation c6129d73-60b1-4f3a-a232-df6e7c4f939f · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

Voice Memory for Agentic Speech Recognition TextGrad: Automatic "Differentiation" via Text

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.392157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.392157Z digest=sha256:dd6a182ee1a38e39307df63fe6f5b3088d03e68cde447a0108c3eb4463f4fd80

Observation 59929868-c901-47ac-a894-95058e5566d8 · outbound

This paper cites strictly greater.

Voice Memory for Agentic Speech Recognition strictly greater

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-01T16:37:24.396537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.396537Z digest=sha256:87dca893da621f32c89e66b317314d4cc02ebf23b66047178537646375ea9454

Pith citing papers

No inbound Pith citation observations are available.