Pith. sign in

Paper Citation Record · LEDGER

Joint Speech Recognition and Speaker Diarization via Sequence Transduction

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 4 inbound Pith citation observations for arXiv:1907.05337.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1907.05337 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T00:59:00.525481Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:32:27.969261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-25T01:00:09.574967Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df2148e4-efbe-47bf-9f6a-169c1c4e18ba · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:00:09.578155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:6b9684ca0b1d3997d36dec99f9440b0c8904bb740f844ddf7eadf7918f5f82e9

Observation dd1d165d-6cab-4e0b-89f6-5acb1a9381b2 · outbound

This paper cites Problem Formulation and Proposed Solution Many machine learning tasks can be expressed as mapping an input sequence into an output sequence.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Problem Formulation and Proposed Solution Many machine learning tasks can be expressed as mapping an input sequence into an output sequence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.300720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:a55ca3db2537e83ca2a49fc65f64e23f4f41b8f4fa6a1535b257f09709085f30

Observation d88a58d1-b5af-475b-8a69-feb4e749a31c · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.362086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:a95863c75872a5c5bb7d391f71e7f614248b0b652de379a978cfbf52b0d988ad

Observation 7993a0b8-0a36-4f24-a645-d9ff7c842276 · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.240446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:4bfa1d375392b5ff2dd535f85927c0519ac44d248546ce92f689a6e45622460e

Observation 29d16264-0740-4a61-bd05-4d6862a3a615 · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.381972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:147e12a287dd76fb04f8a6ae1417f840fddcb714256cc716b43f1d44ae035782

Observation 97e3ca92-1e81-4f84-bc05-55482e9c14ab · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.228658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:0441c4391c51ca81352007b805437460fbde314a7849d57eaf82b7c876973be0

Observation 98cd67ce-d600-4ad2-b90a-5ac506107060 · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.236456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:c8b526e48aec3dcae5f0649a8012054369e8e63620e25a3d64b7f107c705f8a9

Observation d3f6adbc-08e1-4b4d-8dc7-68c887a5b5a9 · outbound

This paper cites We demonstrated the performance of our approach by evaluating it on a large corpus of clinical conversa- tions between physicians and patients.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction We demonstrated the performance of our approach by evaluating it on a large corpus of clinical conversa- tions between physicians and patients

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.292305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:8515a604baa1ac66ef49d6d89e291e4c64e3c207c1d9193559a645efea3ff5db

Observation 07b88adf-565a-46cb-a228-2d068797be7f · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.366683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:8fb27cb4ca81dda0309646ed563c1eb5bfb91bb6b5414ab3ede3ab8f081f8ad4

Observation 175393fa-3556-4888-ad78-2b18e7e2507a · outbound

This paper cites An overview of auto- matic speaker diarization systems.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction An overview of auto- matic speaker diarization systems

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.375332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:475ade988cb878a5fcff4de0763dd66a24967600582a85d3bfc1e703b84a5864

Observation 5e9e5dc7-7e68-4d98-a67c-e79a4aedee57 · outbound

This paper cites Speaker diarization: A review of recent research.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker diarization: A review of recent research

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.378649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:d4158513253846506da6c8fb19129d998675356fbfcdbb4d1eb880361a747fc0

Observation e92e4095-589b-462b-a11f-c7b73fd89946 · outbound

This paper cites A robust speaker clustering algo- rithm.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction A robust speaker clustering algo- rithm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.350458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:79035cc9274b433f3a8cb80e53efb54a57ba15cb644cce956c2071004a9cbd5e

Observation 23ca0396-8c5c-4579-b7e9-b1e22ac3d0ae · outbound

This paper cites Multistage speaker diarization of broadcast news.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Multistage speaker diarization of broadcast news

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.334842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:d8c93c4b3063e56edefc2fc4d1287598ca138967db9fd23076ccf5e728be4f8c

Observation 8fd2b52d-a994-4177-ba77-294dc10d82d6 · outbound

This paper cites Speaker diarization with PLDA i-vector scoring and unsupervised calibration.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker diarization with PLDA i-vector scoring and unsupervised calibration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.287755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:969582d062eb2990bfcb1920b2b01ccdb10dccac5aa93586a9968c0ded40b508

Observation e487747d-1160-474c-b7ad-dbc58b25a482 · outbound

This paper cites Speaker diarization using deep neural network embeddings.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker diarization using deep neural network embeddings

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.393043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:4e6517568fbdaf4e75b9e7e3b9925b9a9ad8512569682e4ba643c1bbbd2b227f

Observation bea5c21e-eceb-46a1-9788-309e58dfb63e · outbound

This paper cites Speaker diarization with LSTM.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker diarization with LSTM

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.338336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:6a4b50a42fae81182cd13a8e78d8346692e4f4a818b2e77b0a1fd424451ba213

Observation fc8682a6-ce90-476f-9257-60ca9c24f8b0 · outbound

This paper cites X-vectors: Robust DNN embeddings for speaker recogni- tion.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction X-vectors: Robust DNN embeddings for speaker recogni- tion

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.358453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:bf97599d9633b652f1a5dbc728228bc147a908d98629de05073dda4f7170b1dd

Observation 8f3d676e-8def-4d48-82c4-8756d3ccabc1 · outbound

This paper cites Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD chal- lenge.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD chal- lenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.371970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:e3a941ddffd4eda9e56112ba46b884e74fdd1109ab4dfcf0fbfc31953b3e4830

Observation 8805d2e1-9e96-4d0e-b8a6-5d50f4562c44 · outbound

This paper cites Tristounet: Triplet loss for speaker turn embedding.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Tristounet: Triplet loss for speaker turn embedding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.389529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:9e4e2dce5a4b22c52bfe8fc7c977490b4f52e0849b998369a2887c1bbdf30a76

Observation 7d2a9063-ea67-4c88-84bb-a4d65a32540a · outbound

This paper cites Fully Supervised Speaker Diarization.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Fully Supervised Speaker Diarization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:00:09.565790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:856b672c03c9de2c8624b48368f05e3b6a06cb3d9b2bfe5cc648ea54776ab79c

Observation cc1057fd-5d2d-46d6-97cc-be24bc01a83a · outbound

This paper cites Speaker di- arization from speech transcripts.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker di- arization from speech transcripts

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.385332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:09aa14752e084406156f279da94efd15863e969f7f019994d8f14e9b25899132

Observation 7ec0a56b-c577-4a66-be1c-ac97361682bc · outbound

This paper cites Multimodal speaker segmentation and diarization using lexical and acoustic cues via sequence to se- quence neural networks.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Multimodal speaker segmentation and diarization using lexical and acoustic cues via sequence to se- quence neural networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.354194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:0a783ebd3e330aec6077e52b2a2cff95750fd7562ec2fa5788cfaef6873a194e

Observation 42913212-c0ef-4597-b76e-e529b5aecee3 · outbound

This paper cites The use of recurrent neural networks in continuous speech recognition.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction The use of recurrent neural networks in continuous speech recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.273558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:e950f00fba329dd1ca5ade02fba869d9c367f5cf3982755b4323f8425f9037b0

Observation f178c1d3-fdcf-4ecd-b75e-b7e18570297d · outbound

This paper cites Connectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Connectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.277767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:d165849917effa4557320de97a97ba864d7f48e30bcc398668db39073aa0651d

Observation bf21111b-9427-4ff4-b75b-b0132871c77a · outbound

This paper cites Sequence transduction with recurrent neural net- works.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Sequence transduction with recurrent neural net- works

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.282214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:8823633dc6c6566dc58d74d5d0b722c8b7539f1761a51258e9cd96c539732f05

Observation 197d9a67-66ea-442a-80cc-916a7cfb3806 · outbound

This paper cites Speech recognition with deep recurrent neural networks.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speech recognition with deep recurrent neural networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.232528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:1b8ba9b5642d937b06b8c55b369e59b12bb657aae46c203e204a5d45bdf2064a

Observation 6eb82b7c-d5fa-4d52-a9cc-cd65ea5ee047 · outbound

This paper cites Streaming End-to-end Speech Recognition For Mobile Devices.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Streaming End-to-end Speech Recognition For Mobile Devices

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:00:09.571824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:9beb61947a2390870a82ac3b164f65e379e76629510b413421fe5f9849152b31

Observation eacc9491-ce87-4ba4-a3cd-4c130cd3a670 · outbound

This paper cites In- datacenter performance analysis of a tensor processing unit.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction In- datacenter performance analysis of a tensor processing unit

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.342565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:9667c7bfede127393526436ea55b04a95479a6f9b6074f4a1391b1d8729c96a5

Observation c5f8e694-5f06-469f-a95c-4dae748b3ed5 · outbound

This paper cites Improving the efficiency of forward-backward algorithm using batched computation in tensorflow.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Improving the efficiency of forward-backward algorithm using batched computation in tensorflow

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.327322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:8ec1fa27d1cf761111ead3d46ea4135e571e308108b7b33bf24da2216a7909f2

Observation 04feb0f1-01c7-4c3c-b3ac-b1108f51d59b · outbound

This paper cites Efficient implementation of recurrent neu- ral network transducer in tensorflow.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Efficient implementation of recurrent neu- ral network transducer in tensorflow

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.305342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:cdca3ed68b9030dd9ff5717130f56c74aebdc3a62edb282c0d5906a70f099dd1

Observation f513f50c-df4a-451d-af65-857af8e9acf2 · outbound

This paper cites Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.396781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:6ba50321928ca97ddbe38c4b1a839b55e8c4619a2ce39460b08b899e12c04d48

Observation 9ec9e717-29f6-49a8-94b9-7bdfed362b15 · outbound

This paper cites Morfessor 2.0: Python implementation and extensions for morfessor base- line.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Morfessor 2.0: Python implementation and extensions for morfessor base- line

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.323290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:47543757a5a4dfec2fbe8b2ff8aa54f1c7be85e922a5d34af3214bf1ce541d69

Observation ddc453d7-d7fe-4be0-8c52-ea030f679765 · outbound

This paper cites Phoneme recognition using time-delay neural networks.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Phoneme recognition using time-delay neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.296298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:a464c362f88c73862b759f1b61e4ca6216e2657f469a1e45751a4f2b88d95d02

Observation 77f9fa79-f727-49da-9c24-555afded2daa · outbound

This paper cites Reducing the computational complexity for whole word models.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Reducing the computational complexity for whole word models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.331126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:dece0d01493269df1ec79e833b582ec4084df145ee98960053379ee12b59352a

Observation b0b87f38-9eb6-403b-9f69-ff41a168d1d7 · outbound

This paper cites Long short-term memory.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Long short-term memory

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.346710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:a2d0abc825c1e3cff855b42c62d445c97b8437a3422ee0a1395abace0680c68e

Observation 86aae55a-6dbc-4151-b93a-d76934d4abf6 · outbound

This paper cites Adam: A method for stochastic opti- mization.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Adam: A method for stochastic opti- mization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.268808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:02db35e95c463e5c7644e38a7fbf5736003ab9e6f3f06a2c8ccbc428d8111700

Observation 9cbfc75a-8ef7-479f-9943-34efdec55b15 · outbound

This paper cites The Rich Transcription Fall 2003 (RT-03F) Evalu- ation Plan.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction The Rich Transcription Fall 2003 (RT-03F) Evalu- ation Plan

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.244402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:bf96b8bd365ec837c6ce5700df814e7104969037212b69af3ca836af4de2ca88

Observation 345aaecc-4234-4cd8-b7af-20877cdb534e · outbound

This paper cites Feature learn- ing with raw-waveform cldnns for voice activity detection.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Feature learn- ing with raw-waveform cldnns for voice activity detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.310789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:81b77d0d7d96384d7a7f26b04fe2902096a9d8a01acd7637efe667d24981610d

Observation 2d39334c-f8b3-4a0e-befd-0d94e98cd608 · outbound

This paper cites End-to- end text-dependent speaker verification.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction End-to- end text-dependent speaker verification

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.314975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:8e631204f7bdd1260bc133483609a034b193185b21d2410deff94c49e4d21e0c

Observation 7f4a7fb6-5a2c-4945-8ac6-00b760f5b051 · outbound

This paper cites V oxceleb2: Deep speaker recognition.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction V oxceleb2: Deep speaker recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.318901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:c81f5149faa98c21c04858e7f769423b72521c2ae243d822ed78293d50d15b3e

Pith citing papers

Observation df2148e4-efbe-47bf-9f6a-169c1c4e18ba · inbound

Joint Speech Recognition and Speaker Diarization via Sequence Transduction cites this paper.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:00:09.578155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:6b9684ca0b1d3997d36dec99f9440b0c8904bb740f844ddf7eadf7918f5f82e9

Observation d9b4c47a-ad4d-405a-b808-72a72c24843f · inbound

Joint ASR and Speaker Role Tagging with Serialized Output Training cites this paper.

Joint ASR and Speaker Role Tagging with Serialized Output Training Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:27.969261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:27.969261Z digest=sha256:0f34cfc6ecab47f53f3cd2505ad8a5a2e89523eb19b1b1c53ed4d8bd29771196

Observation 80d5c565-14c0-4f84-9a26-578fc0102d5e · inbound

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models cites this paper.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.818718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.818718Z digest=sha256:6a1439e2c16c3943834c7f7c041a1ba8dcb435335a7b46d531cfb026ab32e27e

Observation 612daea9-4d17-4202-b167-cc2aa4c16239 · inbound

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation cites this paper.

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:27:29.255358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:27:29.255358Z digest=sha256:83a3609bd08aaa267dc09fc55eaebc0c861d82de51622111b7ccd1c76f873f34