Pith. sign in

Paper Citation Record · LEDGER

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2505.20564.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20564 v3

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:29.985902Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:25.830154Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 11c1482e-737f-49a5-b229-e89af5940eed · outbound

This paper cites While notable progress has been made in speech processing, African languages – including our focus languages, Igbo, Hausa, and Yoruba – have largely been left behind [ 7, 8].

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages While notable progress has been made in speech processing, African languages – including our focus languages, Igbo, Hausa, and Yoruba – have largely been left behind [ 7, 8]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:37.383518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:25.773223Z digest=sha256:f882e347cfbcae15ee5fd30e074da7d977e366b5a8ee31f9a76d93558fa424ce

Observation 3d6feb57-e0c6-4767-b7e1-11026b590d00 · outbound

This paper cites The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:25.830154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:25.830154Z digest=sha256:077c6cff1b94b47e8e8bc2859d90573688d99bde4ec2827132b42ea689cc8f76

Observation 9af524c3-5de4-4e79-9692-67bc1d436ff8 · outbound

This paper cites It features a wide range of speech patterns influenced by age, education levels, accents, and speaking styles – from broken to formal speech, ethnic and dialectal influences.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages It features a wide range of speech patterns influenced by age, education levels, accents, and speaking styles – from broken to formal speech, ethnic and dialectal influences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:37.166709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:25.900230Z digest=sha256:d8ae7f632d4acddf0176d2512f152c56e5f274890feda4cb88e2c1c9d75e6ef6

Observation 8ed4b76b-fcba-4576-a741-80f39a82d5fa · outbound

This paper cites Concretely, we finetune three selected ASR models on our dataset and evaluate them on both our test set (NV Test) and the FLEURS test set [20].

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Concretely, we finetune three selected ASR models on our dataset and evaluate them on both our test set (NV Test) and the FLEURS test set [20]

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:56:30.806189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:25.988475Z digest=sha256:41ee7d467c50c3b5c4b2f7c994bff4986087838940b5899145c71fe134b0eb2e

Observation 203c6a02-ac4b-4ebf-bff6-8e7b517a217f · outbound

This paper cites Built on the principles of ‘data farming’, our approach fosters a symbiotic relationship with language communities.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Built on the principles of ‘data farming’, our approach fosters a symbiotic relationship with language communities

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.866653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.055703Z digest=sha256:2441644b55a58898a1de3871c9c21fe06d0fb8193b11c2efafd8a2d3102d6f51

Observation 5db42c02-a711-4017-bd61-39da9762e4ee · outbound

This paper cites an unresolved cited work.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:36.719269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.126261Z digest=sha256:5df21609016df4c9160ed2cbe2e26dba4cb57cc6a2a383d11ef53e5521144013

Observation bb12c951-1bbc-4dd9-9b29-8f898acf4929 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Robust speech recognition via large-scale weak supervision,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.513465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.179296Z digest=sha256:ab0a1bce9b4d6753c4d6bedff1debc778e50f19ca013cd1fb6bcc119ed1636ba

Observation f421c7e7-1285-4a25-b393-dc14c5d0898b · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Scaling speech technology to 1,000+ languages,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.276579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.261837Z digest=sha256:1e5c58ff368f005617a2777ac53f00f101aed98dfb83824729449042b4aa6518

Observation 83d1bb78-5b5f-407a-81cd-9c9a17bc186b · outbound

This paper cites The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.333107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.333107Z digest=sha256:ddc0de969db7ea493f500e11765ab08e0af38d861c868c471002d6668642f2c6

Observation 44c787d9-56e3-4010-9eee-db697332afbe · outbound

This paper cites IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.034575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.416637Z digest=sha256:bf59e187d3f4ce17782ff5377ddb06ee955f33e3880ca38724b89a5870ced4ae

Observation e1b1d2cb-26c7-4fae-b3ce-93280571f454 · outbound

This paper cites IndicV oices-R: Unlocking a massive multilingual multi- speaker speech corpus for scaling indian TTS,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages IndicV oices-R: Unlocking a massive multilingual multi- speaker speech corpus for scaling indian TTS,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.846060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.464410Z digest=sha256:e132c4e885b4326e5ccbef2be3bf70d74eebd7e35aa047c46d829ba0a01bff19

Observation 89b175f2-8933-40e3-a822-0401697b2d10 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.515597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.515597Z digest=sha256:7a6ba3d938ab292bdfd90f7c4d685a103cc87f2d5fefc9037aef5853c197ab7c

Observation b8e8a9e4-8227-4f6c-a7af-ae01145eebfa · outbound

This paper cites Replication data for Igbo Natural Language Processing Tasks I,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Replication data for Igbo Natural Language Processing Tasks I,

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T13:56:30.348444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.604471Z digest=sha256:9ecf63acb5f015cdaf9c96a8afc5a8b18ac6fb3e189a387b3d717c26e340533f

Observation bdfb1d8e-de7e-46a0-b35f-567ccf5dbe1a · outbound

This paper cites Multi- lingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Multi- lingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.621338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.691212Z digest=sha256:b3c47bdf5c4e59a25bf8c5ef41fbeec17c302142c48264050c32bf3efe2ec1ae

Observation 9830b3b2-4efb-4e9c-95c7-851e53957cb6 · outbound

This paper cites Masakhane -- Machine Translation For Africa.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Masakhane -- Machine Translation For Africa

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.796632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.796632Z digest=sha256:6454beed2b70164484f747aaceb75bb01a325724b8ff5ae74e1ac5c298381d28

Observation a4114642-5a1d-424e-89c3-e0f5d75afd6b · outbound

This paper cites Partici- patory research for low-resourced machine translation: A case study in African languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Partici- patory research for low-resourced machine translation: A case study in African languages,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.407955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:26.930277Z digest=sha256:2ea2984b94f67231687e43443c1963106d20345ef0e8959851ee1e0195aac674

Observation 4ad9046d-5454-4b94-961f-eb494a51f6a6 · outbound

This paper cites The state and fate of linguistic diversity and inclusion in the NLP world,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The state and fate of linguistic diversity and inclusion in the NLP world,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.176512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:27.058131Z digest=sha256:1ce026480d21dbc715024ca2fc8a07c1eae9f53439d646f9aec12fd4e6639500

Observation e8546569-57eb-4caa-8ebf-9c30d18b0c80 · outbound

This paper cites A few thousand trans- lations go a long way! Leveraging pre-trained models for African news translation,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages A few thousand trans- lations go a long way! Leveraging pre-trained models for African news translation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.026735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:27.228195Z digest=sha256:f59d7d924e00289821a8d722640a68e6cc1905ea7d930a03e5c623f0b99e3928

Observation cffa84ed-3bc5-4ca7-bc9d-1a9b04888149 · outbound

This paper cites an unresolved cited work.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:34.803033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:27.343643Z digest=sha256:7bffb7929fe22b55836584d81e3b4af188d525ae8618049e0eb9d9814995bbcd

Observation f9a5659e-f323-45da-ad3a-76927b2cf504 · outbound

This paper cites GlobalPhone: A multi- lingual text & speech database in 20 languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages GlobalPhone: A multi- lingual text & speech database in 20 languages,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.531666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:27.466576Z digest=sha256:d24730141490b7197cd8b5cbd7ec49b6ffdb0e90a36b16ac248c4d2628e8f296

Observation 99dd0dba-51ca-4205-97d5-99d22f6feba7 · outbound

This paper cites YFACC: A Yor`ub´a speech–image dataset for cross-lingual keyword localisation through visual grounding,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages YFACC: A Yor`ub´a speech–image dataset for cross-lingual keyword localisation through visual grounding,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.335515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:27.610794Z digest=sha256:75c61a1a08bad7c8fa8b4bf01ba18ff2c0cd725dc77aa67fbd7c6d961648c9c9

Observation 9c46c67d-597e-4272-8e55-8bb5173a78a1 · outbound

This paper cites \`{I}r\`{o}y\`{i}nSpeech: A multi-purpose Yor\`{u}b\'{a} Speech Corpus.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages \`{I}r\`{o}y\`{i}nSpeech: A multi-purpose Yor\`{u}b\'{a} Speech Corpus

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:56:30.593979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:27.728384Z digest=sha256:a192fbc42000d7dde5be958269c186352d826e7dc816eed3b77bbd4b76f67c1f

Observation 305af219-7afc-4346-a0dd-8610269909f9 · outbound

This paper cites V oices Unheard: NLP resources and models for Yor`ub´a regional dialects,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages V oices Unheard: NLP resources and models for Yor`ub´a regional dialects,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.128670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:27.834694Z digest=sha256:68f7c4f1c526cb5398c949f3e017af8cc30f9944f8933db1c5b2f4988e3e2673

Observation 0a70f040-138c-4672-aea9-90a2801964fa · outbound

This paper cites BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.929958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:27.942667Z digest=sha256:f7e4e328648eaf16b978701681a9eda3cc3902714b231e6cd6e40d5607b19b31

Observation 3147b531-225a-4bb9-ace9-2a0d06c27f34 · outbound

This paper cites Common V oice: A massively-multilingual speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Common V oice: A massively-multilingual speech corpus,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.713462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:28.035532Z digest=sha256:2bf3af18c0c6b9e1671040df4c508df4f9fad2e21e1d0057d3f80c5b16fdc600

Observation a7193215-c851-4cc5-ba8c-f30b9d110db3 · outbound

This paper cites FLEURS: Few-shot learning evaluation of universal representations of speech,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages FLEURS: Few-shot learning evaluation of universal representations of speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.504591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:28.199221Z digest=sha256:6daadf1072d46ffef7825acd3b0546049655d93d258e2381f3911eb4007a146a

Observation cc6f6625-dab4-47ea-9537-5f679ec36301 · outbound

This paper cites Quality at a glance: An audit of web-crawled multilingual datasets,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Quality at a glance: An audit of web-crawled multilingual datasets,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.338427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:28.332611Z digest=sha256:68c3eb11c0bd64dfdcbf8bc7d5201cb99605ac25167c670b7cec38e49e66c2c8

Observation 96873929-315a-446b-b9c0-09bbd1e25e6c · outbound

This paper cites Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced African languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced African languages,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.149536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:28.490123Z digest=sha256:16013ac585058034f8e6699d4d9c3158dc9c519c8d0c227921cb181748c24a41

Observation 8bc08eac-f04b-4a3f-8b00-637b31757013 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:28.580038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:28.580038Z digest=sha256:d63b764b312922c13993cf97cdfabf9459bfbfeac2bde26a8646427f4610cea7

Observation acc0dc36-868a-438c-abd2-87d3171859db · outbound

This paper cites JW300: A wide-coverage parallel corpus for low-resource languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages JW300: A wide-coverage parallel corpus for low-resource languages,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.979742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:28.668877Z digest=sha256:8f14edc4bdf52ab15eb59684f305ffbdc47599a328817076474ccecbdea04553

Observation b7d45a26-68e8-4323-a7b1-ff71d126d160 · outbound

This paper cites `Ir`oy`ınspeech: Yor`ub´a speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages `Ir`oy`ınspeech: Yor`ub´a speech corpus,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.809190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:28.747396Z digest=sha256:785aa999e418b1b3ac698f34c0434ff530ad81c0d8a651387dd8aff0fff7a1c1

Observation 26073f54-6068-493c-ac4a-fe7e5f0122ef · outbound

This paper cites Hausa speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Hausa speech corpus,

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T13:56:30.221593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:28.816780Z digest=sha256:35e389f6afbd090df85fb5d2cb870f8920e40bc25ca00d1f6daaded6e21d62e0

Observation c7af0928-5d8e-4744-88fe-e08c336cb797 · outbound

This paper cites Kencorpus: A Kenyan language corpus of Swahili, Dholuo and Luhya for natural language processing tasks,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Kencorpus: A Kenyan language corpus of Swahili, Dholuo and Luhya for natural language processing tasks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.526703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:28.892248Z digest=sha256:db7c082587ac8eaa1bd26e5672e5e41c81e426eff6b9863249548e59ce26eec7

Observation 2598dc06-a744-406e-8ddd-f238606782fa · outbound

This paper cites LIII. On lines and planes of closest fit to systems of points in space,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages LIII. On lines and planes of closest fit to systems of points in space,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.254430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:29.019144Z digest=sha256:77e6589ffa35ddb226fff1d22c94c79ef4ba68e0355f68f5fd945e21fb3a188f

Observation cd5e4632-baf8-4fb5-b79a-656297daf12e · outbound

This paper cites Robust signal-to-noise ratio estima- tion based on waveform amplitude distribution analysis,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Robust signal-to-noise ratio estima- tion based on waveform amplitude distribution analysis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.952490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:29.122151Z digest=sha256:369895be29db15b0b7a9ede26a49c71783bb2fc0c44d4df12319e7cc47461c67

Observation 5cca4861-c392-4c9a-b371-1c33e94ae017 · outbound

This paper cites Audio quality feature,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Audio quality feature,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.618588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:29.220939Z digest=sha256:6cafe0c9b2be29a896cd5f036b19f23858604608842b575f632522498f675db8

Observation a364f504-22f4-479d-bc69-db6c22758a11 · outbound

This paper cites (n.d.) Evaluation.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages (n.d.) Evaluation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.388331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:29.281822Z digest=sha256:809f302f141d999568161e130532f1cda08a68bf22e0fc101360762391376125

Observation 8f7bbb81-62db-4296-bed6-c80cd69f1d15 · outbound

This paper cites Unsupervised cross-lingual representation learning for speech recognition,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unsupervised cross-lingual representation learning for speech recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.170421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:29.372393Z digest=sha256:00dc4ebda3f9e1415a7349e3363f7ba3bece5e8e4d8354d1eac035f98d85db3c

Observation 6dbcbbbd-4370-4a97-9aba-c82c3a6a2db4 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:29.523320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:29.523320Z digest=sha256:f20cb59237c29a94651d5807f1854299e1d2b586d7a9298c73be5d1aeda606f6

Observation b289d64b-ca04-497b-936c-3f7973f29acd · outbound

This paper cites Small Data? No Problem! Ex- ploring the viability of pretrained multilingual language mod- els for low-resourced languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Small Data? No Problem! Ex- ploring the viability of pretrained multilingual language mod- els for low-resourced languages,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.052468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:29.620364Z digest=sha256:fec985e1defc1fa9f3d0264efb127b93008c3c55dca3c3593a6d04bb95010b59

Observation afc908dd-cfef-403e-8a61-5063e67c5c35 · outbound

This paper cites Data Collection and Quality Challenges in Deep Learning: A Data-Centric AI Perspective.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Data Collection and Quality Challenges in Deep Learning: A Data-Centric AI Perspective

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:29.695164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:29.695164Z digest=sha256:e650471b7923eea9c20cfc62aaa16bb840a0a76fc1c0f410c54fbdac60a838fb

Observation 1adb9986-0c8f-4d73-b305-abffaafe2324 · outbound

This paper cites What makes a high-quality training dataset for large language mod- els: A practitioners’ perspective,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages What makes a high-quality training dataset for large language mod- els: A practitioners’ perspective,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:30.929427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:29.985902Z digest=sha256:af969009d99797b96f1464418a9d47f9298c77955092ba1843ff88f68288b4cd

Observation 88dd9e44-4744-41c2-b976-6488a62e9e9c · outbound

This paper cites A Proposal to Study "Is High Quality Data All We Need?".

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages A Proposal to Study "Is High Quality Data All We Need?"

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:56:30.468477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:56:29.888274Z digest=sha256:3bb2e742d2b506aa3026b83b29f602e78e65065d8dd4ed2a0fde0a15d246fc67

Pith citing papers

Observation 3d6feb57-e0c6-4767-b7e1-11026b590d00 · inbound

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages cites this paper.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:25.830154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:25.830154Z digest=sha256:077c6cff1b94b47e8e8bc2859d90573688d99bde4ec2827132b42ea689cc8f76

Observation 6771a249-08d4-4e15-a0d4-0f90b1c86d08 · inbound

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI cites this paper.

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:50:28.342515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:27:18.774649Z digest=sha256:48bd70c335f96e88e9ea755a78d3a08700516e744a3cbee2af7522be1c8c83ec

Observation 983a31b4-e4ab-4f69-af0b-3058fad203f5 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.629765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:fc6e5cc26763d69b6ccf8fccf52a13966d23dcf844c214a8811e5308273d0c07