Pith. sign in

Paper Citation Record · LEDGER

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2505.20564.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20564 v3

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:29.985902Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:25.830154Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 11c1482e-737f-49a5-b229-e89af5940eed · outbound

This paper cites While notable progress has been made in speech processing, African languages – including our focus languages, Igbo, Hausa, and Yoruba – have largely been left behind [ 7, 8].

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages While notable progress has been made in speech processing, African languages – including our focus languages, Igbo, Hausa, and Yoruba – have largely been left behind [ 7, 8]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:37.383518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:25.773223Z digest=sha256:627f9ef89d4c4df88c2e046db072ae0c6f46f655db491cecbb4d218d41e91a8f

Observation 3d6feb57-e0c6-4767-b7e1-11026b590d00 · outbound

This paper cites The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:25.830154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:25.830154Z digest=sha256:71ba1a61b2ee2208f385fe9791a1a2de059ecfc320357188a642c1f502606d93

Observation 9af524c3-5de4-4e79-9692-67bc1d436ff8 · outbound

This paper cites It features a wide range of speech patterns influenced by age, education levels, accents, and speaking styles – from broken to formal speech, ethnic and dialectal influences.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages It features a wide range of speech patterns influenced by age, education levels, accents, and speaking styles – from broken to formal speech, ethnic and dialectal influences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:37.166709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:25.900230Z digest=sha256:5576dd18d70094d2d85551f7bff70fe865425db1c17c2a85083be15ee9ca2e04

Observation 8ed4b76b-fcba-4576-a741-80f39a82d5fa · outbound

This paper cites Concretely, we finetune three selected ASR models on our dataset and evaluate them on both our test set (NV Test) and the FLEURS test set [20].

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Concretely, we finetune three selected ASR models on our dataset and evaluate them on both our test set (NV Test) and the FLEURS test set [20]

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:56:30.806189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:25.988475Z digest=sha256:57765583922b47705b4b1fe719ef14bfbb03641a3600e9d5a7eb6561cdc9ad45

Observation 203c6a02-ac4b-4ebf-bff6-8e7b517a217f · outbound

This paper cites Built on the principles of ‘data farming’, our approach fosters a symbiotic relationship with language communities.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Built on the principles of ‘data farming’, our approach fosters a symbiotic relationship with language communities

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.866653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.055703Z digest=sha256:8ca941a0d41e1d0581235ea00343f39bfe0a42ab69422e29c4d7ed84bd0c893d

Observation 5db42c02-a711-4017-bd61-39da9762e4ee · outbound

This paper cites an unresolved cited work.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:36.719269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.126261Z digest=sha256:446d95e30e6b4646a66dc3b4415cbff6a3152a09d316cc1b9b5a721b7ac43f29

Observation bb12c951-1bbc-4dd9-9b29-8f898acf4929 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Robust speech recognition via large-scale weak supervision,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.513465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.179296Z digest=sha256:b3eca7ceab049f3fd598977a40f87673481f60b97fc322294c1072a52cf9cf1b

Observation f421c7e7-1285-4a25-b393-dc14c5d0898b · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Scaling speech technology to 1,000+ languages,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.276579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.261837Z digest=sha256:00c1cb10392f768da569506567663227b1c1fa6c50cf6a1925c6898126bf2c72

Observation 83d1bb78-5b5f-407a-81cd-9c9a17bc186b · outbound

This paper cites The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.333107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.333107Z digest=sha256:7c662289f8a6566fbee19e5c3999005b2f3fe3e6eed6c8c71b0bd321c839b26a

Observation 44c787d9-56e3-4010-9eee-db697332afbe · outbound

This paper cites IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.034575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.416637Z digest=sha256:e52f239a7cf0d7c6367d16af487c264ea0ef7171f3a458a3d60c9b336c368983

Observation e1b1d2cb-26c7-4fae-b3ce-93280571f454 · outbound

This paper cites IndicV oices-R: Unlocking a massive multilingual multi- speaker speech corpus for scaling indian TTS,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages IndicV oices-R: Unlocking a massive multilingual multi- speaker speech corpus for scaling indian TTS,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.846060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.464410Z digest=sha256:e63fe6787fb4b1e11d0fc97255d202a03dcfe67be375e7f86eb2bee06d573f37

Observation 89b175f2-8933-40e3-a822-0401697b2d10 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.515597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.515597Z digest=sha256:57f8062628acfffd5faaf121131fd9337f2d8e5c43a786bb99a64ad01375e1bd

Observation b8e8a9e4-8227-4f6c-a7af-ae01145eebfa · outbound

This paper cites Replication data for Igbo Natural Language Processing Tasks I,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Replication data for Igbo Natural Language Processing Tasks I,

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T13:56:30.348444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.604471Z digest=sha256:96225616ba7b6c18374cb7ac07668d1e70d715372992fbe997e4f38e0a24a20c

Observation bdfb1d8e-de7e-46a0-b35f-567ccf5dbe1a · outbound

This paper cites Multi- lingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Multi- lingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.621338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.691212Z digest=sha256:f514aa2d0b6dd0ff9b3950ba6e4c650fdbc322d168c31eab46c0c7c01f8bc2bd

Observation 9830b3b2-4efb-4e9c-95c7-851e53957cb6 · outbound

This paper cites Masakhane -- Machine Translation For Africa.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Masakhane -- Machine Translation For Africa

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.796632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.796632Z digest=sha256:6a3b7172ef82fb17082c2ae7171b7a2267e1cbd30c575a80007f9b0af558ac7a

Observation a4114642-5a1d-424e-89c3-e0f5d75afd6b · outbound

This paper cites Partici- patory research for low-resourced machine translation: A case study in African languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Partici- patory research for low-resourced machine translation: A case study in African languages,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.407955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:26.930277Z digest=sha256:b0a00625844c3a00c3251b88104763bd967159e79dfe146acb768c53b938632f

Observation 4ad9046d-5454-4b94-961f-eb494a51f6a6 · outbound

This paper cites The state and fate of linguistic diversity and inclusion in the NLP world,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The state and fate of linguistic diversity and inclusion in the NLP world,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.176512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:27.058131Z digest=sha256:ecebc30aec3c050e3a6f0c9c976d8e62c5985bb8d19e2a33cf31136ea2b35de5

Observation e8546569-57eb-4caa-8ebf-9c30d18b0c80 · outbound

This paper cites A few thousand trans- lations go a long way! Leveraging pre-trained models for African news translation,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages A few thousand trans- lations go a long way! Leveraging pre-trained models for African news translation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.026735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:27.228195Z digest=sha256:3402cbb9ab2e03aadfe15dbf0caa2ea9043a3c7ddd2225041397d8637d08c6fc

Observation cffa84ed-3bc5-4ca7-bc9d-1a9b04888149 · outbound

This paper cites an unresolved cited work.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:34.803033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:27.343643Z digest=sha256:5fc531846bcdffaf7802d21e7df25b034bf49a4dee2abeedc762a5b657e58243

Observation f9a5659e-f323-45da-ad3a-76927b2cf504 · outbound

This paper cites GlobalPhone: A multi- lingual text & speech database in 20 languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages GlobalPhone: A multi- lingual text & speech database in 20 languages,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.531666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:27.466576Z digest=sha256:43c987bc366626fa9dc6ca60ac0cb2a70f02e173138ea9a4f6958035c5e5b00d

Observation 99dd0dba-51ca-4205-97d5-99d22f6feba7 · outbound

This paper cites YFACC: A Yor`ub´a speech–image dataset for cross-lingual keyword localisation through visual grounding,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages YFACC: A Yor`ub´a speech–image dataset for cross-lingual keyword localisation through visual grounding,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.335515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:27.610794Z digest=sha256:d6ea70511d9a914f610ff952ae9b76a138ff7f461d6afea232f3955baf82f929

Observation 9c46c67d-597e-4272-8e55-8bb5173a78a1 · outbound

This paper cites \`{I}r\`{o}y\`{i}nSpeech: A multi-purpose Yor\`{u}b\'{a} Speech Corpus.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages \`{I}r\`{o}y\`{i}nSpeech: A multi-purpose Yor\`{u}b\'{a} Speech Corpus

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:56:30.593979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:27.728384Z digest=sha256:228ccb8696a2cf81ca6737308bbb8c7243d462d37e4f333f620362ae7ff69c16

Observation 305af219-7afc-4346-a0dd-8610269909f9 · outbound

This paper cites V oices Unheard: NLP resources and models for Yor`ub´a regional dialects,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages V oices Unheard: NLP resources and models for Yor`ub´a regional dialects,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.128670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:27.834694Z digest=sha256:6e2313c675ce2430314acc5b9428714af1bc687d8a8128ec92396aed1626ed2e

Observation 0a70f040-138c-4672-aea9-90a2801964fa · outbound

This paper cites BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.929958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:27.942667Z digest=sha256:8fc0c3407aca0c1d7a4cfedacf5edb3fc1e8f5b854988bb0181c64e109dd6358

Observation 3147b531-225a-4bb9-ace9-2a0d06c27f34 · outbound

This paper cites Common V oice: A massively-multilingual speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Common V oice: A massively-multilingual speech corpus,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.713462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:28.035532Z digest=sha256:7287fbe4d41b501002ff16bf2a222756c287979a1dd2cb336efff46228ec9826

Observation a7193215-c851-4cc5-ba8c-f30b9d110db3 · outbound

This paper cites FLEURS: Few-shot learning evaluation of universal representations of speech,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages FLEURS: Few-shot learning evaluation of universal representations of speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.504591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:28.199221Z digest=sha256:f3ac36a6453482b02c06de50c1f3a87e83614bbbf8cb1aacd21f97a4441cb6b6

Observation cc6f6625-dab4-47ea-9537-5f679ec36301 · outbound

This paper cites Quality at a glance: An audit of web-crawled multilingual datasets,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Quality at a glance: An audit of web-crawled multilingual datasets,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.338427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:28.332611Z digest=sha256:ceacc5ad10f6c29bc2a92975a53fc26903a24f44a92ad9b4e065ff1143820b7d

Observation 96873929-315a-446b-b9c0-09bbd1e25e6c · outbound

This paper cites Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced African languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced African languages,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.149536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:28.490123Z digest=sha256:70e4257a5af92bdaf6f1ca3313936f72da408dfce01cf0215af3c413e5cba39e

Observation 8bc08eac-f04b-4a3f-8b00-637b31757013 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:28.580038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:28.580038Z digest=sha256:7ff64f959f84349e4317e47ef08da0152afbfeb2c292d77613e26f731f63297c

Observation acc0dc36-868a-438c-abd2-87d3171859db · outbound

This paper cites JW300: A wide-coverage parallel corpus for low-resource languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages JW300: A wide-coverage parallel corpus for low-resource languages,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.979742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:28.668877Z digest=sha256:75d3c6c3ef99b64b1435da6372a6b89aead5aee58928c39cff02ea035d4e0ab6

Observation b7d45a26-68e8-4323-a7b1-ff71d126d160 · outbound

This paper cites `Ir`oy`ınspeech: Yor`ub´a speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages `Ir`oy`ınspeech: Yor`ub´a speech corpus,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.809190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:28.747396Z digest=sha256:fca21962ed7c186861907450f8f4ab8bba6ea08d89d21d396cbe93b67c9c90ab

Observation 26073f54-6068-493c-ac4a-fe7e5f0122ef · outbound

This paper cites Hausa speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Hausa speech corpus,

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T13:56:30.221593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:28.816780Z digest=sha256:cf96a6304d76ef6b9bbeae685b6050dba547adfc035aad7f37161fbf1e90f727

Observation c7af0928-5d8e-4744-88fe-e08c336cb797 · outbound

This paper cites Kencorpus: A Kenyan language corpus of Swahili, Dholuo and Luhya for natural language processing tasks,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Kencorpus: A Kenyan language corpus of Swahili, Dholuo and Luhya for natural language processing tasks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.526703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:28.892248Z digest=sha256:2704b20c45b7920a99eac841972f2263092e239c31d427bb870a57d4a46d7cce

Observation 2598dc06-a744-406e-8ddd-f238606782fa · outbound

This paper cites LIII. On lines and planes of closest fit to systems of points in space,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages LIII. On lines and planes of closest fit to systems of points in space,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.254430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:29.019144Z digest=sha256:92d44826fb2d5447e068d87d55962bb88f440e12df1b46e27d79c32f0638d790

Observation cd5e4632-baf8-4fb5-b79a-656297daf12e · outbound

This paper cites Robust signal-to-noise ratio estima- tion based on waveform amplitude distribution analysis,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Robust signal-to-noise ratio estima- tion based on waveform amplitude distribution analysis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.952490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:29.122151Z digest=sha256:1fbd29fec5471fc49130b59c1f791f09842a1ed3d31ad7818c1aa785db221fd1

Observation 5cca4861-c392-4c9a-b371-1c33e94ae017 · outbound

This paper cites Audio quality feature,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Audio quality feature,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.618588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:29.220939Z digest=sha256:b1e2e999f45635f92bcfdfa7f2b2af2d8c1ac11dac7bc62064fca394d55793e4

Observation a364f504-22f4-479d-bc69-db6c22758a11 · outbound

This paper cites (n.d.) Evaluation.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages (n.d.) Evaluation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.388331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:29.281822Z digest=sha256:0a676ee38e0b569b2f95905e0f6588969006049e4727a9dddc020c84c4d3871b

Observation 8f7bbb81-62db-4296-bed6-c80cd69f1d15 · outbound

This paper cites Unsupervised cross-lingual representation learning for speech recognition,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unsupervised cross-lingual representation learning for speech recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.170421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:29.372393Z digest=sha256:5f11942b94a1b83b82672dcfe9f7733f51ba904d936b072305e4fd673b731848

Observation 6dbcbbbd-4370-4a97-9aba-c82c3a6a2db4 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:29.523320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:29.523320Z digest=sha256:bfa619ca02e4e5045f3d27f600f4fa4ac39c07b54e517c8ddc8872c246a69b33

Observation b289d64b-ca04-497b-936c-3f7973f29acd · outbound

This paper cites Small Data? No Problem! Ex- ploring the viability of pretrained multilingual language mod- els for low-resourced languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Small Data? No Problem! Ex- ploring the viability of pretrained multilingual language mod- els for low-resourced languages,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.052468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:29.620364Z digest=sha256:bca010305b71e5eb8b25d126c6e6952ae84deb911ecadb7402376ca639c80664

Observation afc908dd-cfef-403e-8a61-5063e67c5c35 · outbound

This paper cites Data Collection and Quality Challenges in Deep Learning: A Data-Centric AI Perspective.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Data Collection and Quality Challenges in Deep Learning: A Data-Centric AI Perspective

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:29.695164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:29.695164Z digest=sha256:8013c89bbcf19738b55fb06c99a1d9ba2c99837c017bae5d814bbe88cef84e95

Observation 1adb9986-0c8f-4d73-b305-abffaafe2324 · outbound

This paper cites What makes a high-quality training dataset for large language mod- els: A practitioners’ perspective,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages What makes a high-quality training dataset for large language mod- els: A practitioners’ perspective,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:30.929427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:29.985902Z digest=sha256:db1c0d91d932076777bed49845ed34c785093457c2d6fe4c860bd23fd231067c

Observation 88dd9e44-4744-41c2-b976-6488a62e9e9c · outbound

This paper cites A Proposal to Study "Is High Quality Data All We Need?".

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages A Proposal to Study "Is High Quality Data All We Need?"

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:56:30.468477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:56:29.888274Z digest=sha256:1b5ebed5eb9c573c3dd13ba7e38d2a4162e10ff57307f52947018b91f6a5b9ca

Pith citing papers

Observation 3d6feb57-e0c6-4767-b7e1-11026b590d00 · inbound

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages cites this paper.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:25.830154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:25.830154Z digest=sha256:71ba1a61b2ee2208f385fe9791a1a2de059ecfc320357188a642c1f502606d93

Observation 6771a249-08d4-4e15-a0d4-0f90b1c86d08 · inbound

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI cites this paper.

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:50:28.342515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T19:27:18.774649Z digest=sha256:bceba6256b584067370acaf0a2dca7238395b6c8fc8295a4189aa8d76949013d

Observation 983a31b4-e4ab-4f69-af0b-3058fad203f5 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.629765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:5b238ce1a5d72a67755858fd98f4b81f5cb41c2f519596d82e6cf9df1a6ff6d8