Back to featured projects

A little more of the work

Paza: Speech Data for Low-Resource African Languages

Research support for speech-data collection, quality control, and ASR evaluation where suitable language datasets are scarce.

Liz's Library
Back to research

Research initiative

Inside the folder

AI Researcher, Internal and Part Time

Microsoft Africa Research Institute (MARI) · February 2026 - Present

Speech AIASRDataset designAfrican languages

Description

Research support for speech-data collection, quality control, and ASR evaluation where suitable language datasets are scarce.

Ongoing internal, part-time research

The research question

How can speech-data collection and evaluation support African languages that are not well represented in existing datasets?

My contribution

  • Assessed dataset gaps and recruited speakers for the voice dataset initiative
  • Led collection, labeling, and annotation, establishing protocols and quality controls
  • Partnered on PazaBench and supported fine-tuned speech models for African languages

Outcomes & evidence

  • ML-ready speech data and quality-control processes for languages with limited suitable vendor datasets
  • Support for ASR benchmarking and fine-tuned models for Swahili, Dholuo, Kalenjin, Kikuyu, Maasai, and Somali

Approach & methods

  • Started with dataset-gap assessment, then coordinated speaker recruitment, collection, labeling, and annotation.
  • Created collection protocols and quality checks where suitable vendor datasets did not exist.
  • Supported benchmarking across languages and the review of fine-tuned models, considering accuracy, data requirements, and inference efficiency.

Scope & limitations

  • This page describes the research scope and contribution, not unpublished dataset sizes, comparative model scores, or internal experimental findings.
  • No language-specific performance claim is presented as a general guarantee.

Sources & project links

Folder openResearch initiative