Pith. sign in

Paper Citation Record · LEDGER

Unified Learnable 2D Convolutional Feature Extraction for ASR

As of 17 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2509.10031.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.10031 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:20:04.427870Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 109d7505-fcb0-4f70-9f06-f3381a8736ef · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Unified Learnable 2D Convolutional Feature Extraction for ASR Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.004092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.004092Z digest=sha256:abf1e1566762501977164e7bbb565334eee8307df073f6705df06e07be91d01c

Observation 48be08d9-1fb7-4f78-8529-917af10b7aca · outbound

This paper cites Comparison of parametric representations for monosyllabic word recognition in contin- uously spoken sentences,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Comparison of parametric representations for monosyllabic word recognition in contin- uously spoken sentences,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.093839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.093839Z digest=sha256:1757f9a471246bf958ef58a0221c6fdad06ed4f7a3aaecb95533d61d03f8c761

Observation f7950395-786a-4c0e-a1e0-ca042a50ee78 · outbound

This paper cites Gamma- tone features and feature combination for large vocabulary speech recognition,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Gamma- tone features and feature combination for large vocabulary speech recognition,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.158499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.158499Z digest=sha256:03827519d52eb10675905c81618aaea5ae8a0f059b61aa8bece7d1b9528acdb7

Observation e4fd773a-b81d-478a-899c-db430e46468b · outbound

This paper cites Convolu- tional, long short-term memory, fully connected deep neural networks,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Convolu- tional, long short-term memory, fully connected deep neural networks,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.273308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.273308Z digest=sha256:623f7be42bfc0c7d844ec699cf2603d1cb8b44ff5b9dccd94d2be6fc5bf1ea5e

Observation d559b726-da31-4802-b6f4-94b6c6cd3265 · outbound

This paper cites Learning the speech front-end with raw wave- form CLDNNs,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Learning the speech front-end with raw wave- form CLDNNs,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.336165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.336165Z digest=sha256:5f04d5df8df3ac1d09e9a64376c5a947041821e3521dba094a3a971e85e3d877

Observation 32b051fe-d3f6-45d1-9a00-b0866c08dc5b · outbound

This paper cites Learning filterbanks from raw speech for phone recognition,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Learning filterbanks from raw speech for phone recognition,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.395692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.395692Z digest=sha256:0a124e9ffb7dc1f7263e8fe788d4c4700624f59b34cd6d0b6f1989d4e80db20a

Observation 152edbf0-153f-4493-a846-e4879bd61ace · outbound

This paper cites Acoustic modeling of speech waveform based on multi-resolution, neural network signal processing,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Acoustic modeling of speech waveform based on multi-resolution, neural network signal processing,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.480646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.480646Z digest=sha256:39c4f0a592ed388fba1396c25e0fc3d6d2b15094892c07c0b036081c2f152177

Observation ed2e966d-a370-4635-98f8-874524fa3751 · outbound

This paper cites LEAF: A learnable frontend for audio classification,.

Unified Learnable 2D Convolutional Feature Extraction for ASR LEAF: A learnable frontend for audio classification,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.570844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.570844Z digest=sha256:b4d8c3d7d384f237f300f2f37079eba598425ae7036a9b82e9cbade1a3d24dd0

Observation 84f0e200-a973-4ed1-a5f1-5b70c354cf00 · outbound

This paper cites Speaker recognition from raw waveform with SincNet,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Speaker recognition from raw waveform with SincNet,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.604119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.604119Z digest=sha256:ca38855327496b91d113052d0fb24ece39ef18d31469e466b1e17ded8db56a72

Observation d8a62980-9637-4de0-aad0-3de202938d25 · outbound

This paper cites Acoustic mod- eling with deep neural networks using raw time signal for LVCSR,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Acoustic mod- eling with deep neural networks using raw time signal for LVCSR,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.664169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.664169Z digest=sha256:c05c7aceeff9110895c861174148e3de267876525cdb28a73c5d43617d1fb52c

Observation 1affca94-1e91-4aee-9c22-9c9c07cb0b69 · outbound

This paper cites Comparative analysis of the wav2vec 2.0 feature extractor,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Comparative analysis of the wav2vec 2.0 feature extractor,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.751959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.751959Z digest=sha256:57a094b60439468bc1e0fb668757e862f08ede02bd1ac9e64252c1576add2d49

Observation 04ccb861-c785-4e2e-a52d-6873fd67a363 · outbound

This paper cites Perception of speech and sound,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Perception of speech and sound,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.852765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.852765Z digest=sha256:6a212188d91506136459213d00ea224e7877a790587d3ce13510d121a422b41b

Observation fd158a96-6040-4ac8-a193-359a6a599c66 · outbound

This paper cites Perceptual linear predictive (PLP) analysis of speech,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Perceptual linear predictive (PLP) analysis of speech,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.945495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.945495Z digest=sha256:07aaac2861ac35ca46dedd3e590038ca27ea3b2d16d65eb53011b8cf58d2cd4e

Observation b38217d0-6ebe-4d38-9af2-cf4c7f95da9e · outbound

This paper cites Very deep convolutional networks for large-scale image recognition.

Unified Learnable 2D Convolutional Feature Extraction for ASR Very deep convolutional networks for large-scale image recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.009824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.009824Z digest=sha256:7c31b64b4f372cdc10e87bb32e6ce471874d9d9ad067190c9ee9f9a51fe90221

Observation 65984a7a-7b30-418a-92dd-1f4361b48cf1 · outbound

This paper cites Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.092129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.092129Z digest=sha256:4b4bbb3cf2d4b8cd794511fd516afce170a2edc2d71462c297179213a53323df

Observation 7a531257-6082-42c8-a61b-4b7232913449 · outbound

This paper cites Very deep convolu- tional networks for end-to-end speech recognition,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Very deep convolu- tional networks for end-to-end speech recognition,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.172520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.172520Z digest=sha256:cf375bf9525c2edf9597115c95aa93d8f23fcd337cde8682e13ce2b5f8a12b2b

Observation 965a78e3-64e3-472c-ac99-b9146e79ea93 · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

Unified Learnable 2D Convolutional Feature Extraction for ASR ESPnet: End-to-end speech processing toolkit,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.255481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.255481Z digest=sha256:610e85f3f1054ff196acdb3a879725160d7292dc9cb4d85bdfd31bc3bedc9353

Observation a63d845f-d60b-4497-80bd-df9a7933dccd · outbound

This paper cites wav2vec: Unsupervised pre-training for speech recogni- tion,.

Unified Learnable 2D Convolutional Feature Extraction for ASR wav2vec: Unsupervised pre-training for speech recogni- tion,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.332929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.332929Z digest=sha256:fdd32368ff659ce29a4438cb7ccfac6ace02e55c2dce9b4c008ce3c7828cbb56

Observation d9a855c8-0096-4f02-96d7-131d64e3b8fa · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech rep- resentations,.

Unified Learnable 2D Convolutional Feature Extraction for ASR wav2vec 2.0: A framework for self-supervised learning of speech rep- resentations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.415639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.415639Z digest=sha256:53b5fb9731cb666fa199b835ff843741ac58b70f7fef725a87e8fcb89305c97b

Observation 1d8b208e-304f-4941-9435-709d9e35762d · outbound

This paper cites On architectures and training for raw waveform feature ex- traction in ASR,.

Unified Learnable 2D Convolutional Feature Extraction for ASR On architectures and training for raw waveform feature ex- traction in ASR,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.495602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.495602Z digest=sha256:7de141f2393733cee84c7838275564709de341c20708f2016e5478c707baae96

Observation cb45b9f4-71f1-486a-a227-605212e575be · outbound

This paper cites HuBERT: How much can a bad teacher bene- fit ASR pre-training?.

Unified Learnable 2D Convolutional Feature Extraction for ASR HuBERT: How much can a bad teacher bene- fit ASR pre-training?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.557499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.557499Z digest=sha256:3bc6e0c53349ee23f222c8ce05712cecd9059edcd795fe88db7098d0852fbaa7

Observation 4d468084-f470-46b5-8eee-afd7be793fb0 · outbound

This paper cites SpecAugment: A simple data aug- mentation method for automatic speech recognition,.

Unified Learnable 2D Convolutional Feature Extraction for ASR SpecAugment: A simple data aug- mentation method for automatic speech recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.635269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.635269Z digest=sha256:2116113b231f0b243e3a6a542bc4b0367a2cca2d37aadba15c134c7138e1eb7d

Observation 62160d4f-c8a6-4ce0-b01e-378c1e907074 · outbound

This paper cites Regularizing learnable feature extraction for auto- matic speech recognition,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Regularizing learnable feature extraction for auto- matic speech recognition,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.736813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.736813Z digest=sha256:d324cfac0dd73ebfb64510dcf1b481ce970737c9d83fdb0f97a55dafd7415d34

Observation 1730b50e-d520-4c9f-ad35-ba72024e14fd · outbound

This paper cites Layer Normalization.

Unified Learnable 2D Convolutional Feature Extraction for ASR Layer Normalization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.773906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.773906Z digest=sha256:eea39ac69c864f51734b32fefe06e9390998bc3138cdc8f77d874bd9e44e6fcf

Observation 9aaa4fa4-7b50-41a4-922b-227e6120239d · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Unified Learnable 2D Convolutional Feature Extraction for ASR Gaussian Error Linear Units (GELUs)

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.847948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.847948Z digest=sha256:1b9e5329f48d69ef1b6a6004c553766b1294e13be3876d05cbd55449aad9fbfd

Observation 9ea69aff-4f94-4b2f-90ee-430ea6c1dc80 · outbound

This paper cites Efficient training of neural transducer for speech recognition,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Efficient training of neural transducer for speech recognition,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:03.926948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:03.926948Z digest=sha256:6dda12131126e32ca8f451a95954c1a28d70b42b1d2cdaef8ac13915facd2255

Observation fd6b6f05-9b20-41ab-9ce4-7b60e3ae3efe · outbound

This paper cites Lib- riSpeech: An ASR corpus based on public domain audio books,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Lib- riSpeech: An ASR corpus based on public domain audio books,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:04.010498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:04.010498Z digest=sha256:47978f86dea1e8c6e6347325ed96aa75623b12640f8c35f5f94c656de2c16132

Observation 83f12331-f177-4233-88a5-a541a6a2ce43 · outbound

This paper cites Joint-sequence models for grapheme- to-phoneme conversion,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Joint-sequence models for grapheme- to-phoneme conversion,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:04.088644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:04.088644Z digest=sha256:e3752fe411b4410a27bd77919a8f8487ec5be83376fe1dce8d9501af6af71f59

Observation b8eae0d9-983d-43d9-ad73-3c0d9fc7b024 · outbound

This paper cites Self-attention with relative position representations,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Self-attention with relative position representations,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:04.146514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:04.146514Z digest=sha256:a3dac1402ede00e11c5a190721d044c4fd652502f1f29ebcd4f0be1e973c1b04

Observation b195cab1-6a15-4cdc-a7fc-272b20d3fb23 · outbound

This paper cites Decoupled weight decay regu- larization,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Decoupled weight decay regu- larization,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:04.228552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:04.228552Z digest=sha256:60edf255c0d0a90e5b644f7dce56fe137587d33d2ce24c3781a1230b04dfb7b1

Observation 82198bf8-362d-4f25-96d0-0f0598006830 · outbound

This paper cites Flashlight: Enabling innovation in tools for machine learning,.

Unified Learnable 2D Convolutional Feature Extraction for ASR Flashlight: Enabling innovation in tools for machine learning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:04.312647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:04.312647Z digest=sha256:cc2ced9b1ada288eba267c0d77786def878a6dbcdb6085087bc19b0b9bfba423

Observation f0e7b3a7-c95b-4a88-b65b-6f8c881cb1c4 · outbound

This paper cites An efficient auditory filterbank based on the gamma- tone function,.

Unified Learnable 2D Convolutional Feature Extraction for ASR An efficient auditory filterbank based on the gamma- tone function,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:04.427870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:04.427870Z digest=sha256:2489d859f0d4a23af6547967acf91e537a94ba6c0ca1303941d2749d424c8324

Pith citing papers

No inbound Pith citation observations are available.