Pith. sign in

Paper Citation Record · LEDGER

Optimizing Speech Multi-View Feature Fusion through Conditional Computation

As of 12 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2501.08057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08057 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:35:13.120441Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88c42a83-1785-410c-8ab9-fb90274451a7 · outbound

This paper cites Introduction to digital speech processing,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Introduction to digital speech processing,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.676357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:12.952741Z digest=sha256:7eeae70f25fd8fa81402795fa33b2daf3d97ed719c18093f875cf50acc0edac6

Observation 1937764e-8328-45c1-90e1-9d91f823195f · outbound

This paper cites Learning robust features using deep learning for automatic seizure detection,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Learning robust features using deep learning for automatic seizure detection,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.664268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:12.957265Z digest=sha256:eea95635050726f73579a24586dfccf73ae5e58b6a7e26511d7777e3db61053f

Observation de5af010-494a-43ef-a812-c3d97c274977 · outbound

This paper cites Listen, Attend and Spell.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Listen, Attend and Spell

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.962138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.962138Z digest=sha256:81bc550cba8ccea97da4fbcf351fe5829b0609fc06521b8603d4b389a567845a

Observation dcc3f86b-bba6-49b1-bb64-2770c4ced7f2 · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.967080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.967080Z digest=sha256:1b685cf859aca58ba13a0205708fee7b6e9fce5f776fe2c5b37dc20b401e5700

Observation 1be994b9-d8e9-43f4-af09-481ad06a238a · outbound

This paper cites Data2vec: A general framework for self-supervised learning in speech, vision and language,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Data2vec: A general framework for self-supervised learning in speech, vision and language,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.652416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:12.972165Z digest=sha256:7d763a302a78eaea9bce66007086d8fff36ec54b25ce07273c98fb7961be18b0

Observation 134a1727-a30e-4c74-9574-32c86366f7ce · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.640340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:12.978136Z digest=sha256:2e67fb5dea681a69601d15cbe990b908d7cfb586de582042341bbf12aee4cd70

Observation 809aaec6-045b-47cc-a99c-3980a4e25ae2 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.628277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:12.983242Z digest=sha256:28c8a2b89299e7462565a0a8b74155657d0fe3353a06d672039088c890d01ff1

Observation b4f68caa-92d2-4541-a491-63382220b722 · outbound

This paper cites Exploration on HuBERT with Multiple Resolu- tions,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Exploration on HuBERT with Multiple Resolu- tions,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.616277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:12.988042Z digest=sha256:c87d9f8fa11256945a5d5af8fe73e87f9bfa1f980ec04800f6ae0451521e9528

Observation e9db1f10-40da-4977-beff-caa965cdb38f · outbound

This paper cites Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.992523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.992523Z digest=sha256:cc6c7597b6fd55ada156e94f44360b2901347b0e4db622b4149b24db14e6c81c

Observation 7b03d0e0-e681-4756-a5ba-ad0797c0c66a · outbound

This paper cites UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.996937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.996937Z digest=sha256:4d2ea1f158b94652864d92599e1e351abb2dd3e5edb7fcbc2db262cacba3da92

Observation 79750105-313c-41cc-81e7-5be56738d061 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.001744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.001744Z digest=sha256:ad8cdcf8b4c2a7186ca5840a7f26daa84a63305e8235b24d9e4b8b3e9abca952

Observation 4ee0998a-0bc6-43dc-9c7c-0a006c6fc662 · outbound

This paper cites Multi- view information-bottleneck representation learning,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Multi- view information-bottleneck representation learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.604733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.007235Z digest=sha256:4f97a554dca76cb3ec17054e2f1ad05646448cc8129eeb776b39ce6c8c62da37

Observation 97adc650-762e-429e-830a-ce6f1a778a6f · outbound

This paper cites Deep multi-view learning methods: A review,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep multi-view learning methods: A review,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.589761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.011623Z digest=sha256:b057239a7ed044090bd9c37d058b3432b71638202b06e45992a1e21952cc5991

Observation 3cf63879-175c-4448-8c8d-431e7a867dbe · outbound

This paper cites The Platonic Representation Hypothesis.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation The Platonic Representation Hypothesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.016293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.016293Z digest=sha256:c501150f222c3d207eac6d4124fe4d7b323d4e4345331825cc8193abde62e54e

Observation 80e901f7-d97f-4080-a2ec-78a4f9ec1370 · outbound

This paper cites Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.574820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.023101Z digest=sha256:86deb8e4bc3ca2487f838a8bfe24da9e2edc3ba24b88d7e84ed6c160ddd590fe

Observation 42892fa8-b830-4431-b653-a619d8cf2d38 · outbound

This paper cites Improving speech emotion recognition by fusing self-supervised learn- ing and spectral features via mixture of experts,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Improving speech emotion recognition by fusing self-supervised learn- ing and spectral features via mixture of experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.559020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.028995Z digest=sha256:4381ea673e5d86814a7469bc4d8ef99b55a0a937d483ddae1b90a055e0ca7b89

Observation 7db3994b-4a62-4b3c-a152-892cfc5cc83b · outbound

This paper cites Deep learning of representations: Looking forward,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep learning of representations: Looking forward,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.544333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.035113Z digest=sha256:f9c391288e600b00b42c72032e578e7e7a166d920289249127f6691c349f4acc

Observation d1f92b65-ee65-4a44-b930-b9bf5021c6f9 · outbound

This paper cites Depth-Adaptive Transformer.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Depth-Adaptive Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.040015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.040015Z digest=sha256:9dd5f0fe863b3153944db7247d9d4146327ed6d0379a210672d5d9e8aa3478b3

Observation 5f92795d-8300-4ff3-a32a-b55de0f7c3a0 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.048303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.048303Z digest=sha256:8e31c8ae791dfa394c2e77987dd9998ce6315b72940d610e33e62be46a2a7da1

Observation 1de2a624-7275-406c-99f7-68507eac3d18 · outbound

This paper cites Modeling task relationships in multi-task learning with multi- gate mixture-of-experts,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Modeling task relationships in multi-task learning with multi- gate mixture-of-experts,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.525266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.054186Z digest=sha256:97f3a85b7de74cf53cabb20da93aa40d200b554d9f9210a18e15964d6af999fe

Observation 263ec31d-b428-4805-8b60-79b344d77f15 · outbound

This paper cites Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.059024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.059024Z digest=sha256:d2e867d87d46c76fd7d769a5aea5d369693e127ee60e59b00609e36e0a6ed70b

Observation fa28073f-17e4-43b5-b200-a7138ab9261d · outbound

This paper cites Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.064469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.064469Z digest=sha256:8bd05878268ba323b28a9fea5f543c953f4e4e82bba173d9e739d12ee789675c

Observation 83c81a68-fcf4-4208-8a2d-4c9dba5f3425 · outbound

This paper cites Soft Alignment of Modality Space for End-to- End Speech Translation,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Soft Alignment of Modality Space for End-to- End Speech Translation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.505664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.069007Z digest=sha256:aec273c479582a5e452052aef74b78b464334a4904bba519f0a6a1b76e8259cf

Observation 0c0b2c8b-78c2-4d45-8c40-a2ae4576bdef · outbound

This paper cites Deep residual learning for image recognition,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep residual learning for image recognition,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.073686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.073686Z digest=sha256:3e4ac6710e22380d262750bab16f5fd3ce4a2b8b887d1b83bfd82f061c831f45

Observation c91f3c40-8959-479e-8e60-55284b580e68 · outbound

This paper cites Progres- sive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Progres- sive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.472636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.078447Z digest=sha256:2ab61a2313fa605ec04b4a7f202f20df7226aab098a775398a08545705989cf6

Observation ae696b85-5208-455c-aa0f-73fe035f49be · outbound

This paper cites Gradient surgery for multi-task learning,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Gradient surgery for multi-task learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.459433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.085525Z digest=sha256:b9474431da0c7c650ae24fb73b386175b0bb1e3380b92212a93b70e9c0918e81

Observation 713c1147-2693-4839-b9f0-74d40a04d96d · outbound

This paper cites Must-c: a multilingual speech translation corpus,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Must-c: a multilingual speech translation corpus,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.445597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.095399Z digest=sha256:423527a184d4926415d36c5821bd94d73412331245fc645d950035541450665c

Observation 71a5f744-1695-4b27-8df8-d5357121a6c1 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Librispeech: an asr corpus based on public domain audio books,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.427050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.101348Z digest=sha256:40fb3dcb758331ad3ced68af2b6bb55ed312849713d03dee76ef0cb18aaeceab

Observation 59eb1700-894c-4115-9694-734437d47877 · outbound

This paper cites Modality Adaption or Regularization? A Case Study on End-to-End Speech Translation.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Modality Adaption or Regularization? A Case Study on End-to-End Speech Translation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:35:13.209587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:35:13.106571Z digest=sha256:0175d935de85b8d26a3302222a1fb6d683acc33ee2209327154bdc1c1cc81fa3

Observation 0b0fa77c-ad3e-4681-a185-8a1832ed8f8f · outbound

This paper cites Attention Is All You Need.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Attention Is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.113877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.113877Z digest=sha256:56ed13d72edac240756d52a5ea7d4ee9f79212b2f0e05f217db27f7026f4255c

Observation 1d095de8-3b52-489a-81f0-adfc5bb1981d · outbound

This paper cites A Call for Clarity in Reporting BLEU Scores.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation A Call for Clarity in Reporting BLEU Scores

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.120441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.120441Z digest=sha256:093e66a58191d919a9a0c22d732e77866ec37b649503cc3544969196ea9867ff

Pith citing papers

No inbound Pith citation observations are available.