Research initiative
Inside the folder
AI Researcher, Internal and Part Time
Microsoft Africa Research Institute (MARI) · February 2026 - Present
Speech AIASRDataset designAfrican languages
Description
Research support for speech-data collection, quality control, and ASR evaluation where suitable language datasets are scarce.
Ongoing internal, part-time research
The research question
How can speech-data collection and evaluation support African languages that are not well represented in existing datasets?
My contribution
- Assessed dataset gaps and recruited speakers for the voice dataset initiative
- Led collection, labeling, and annotation, establishing protocols and quality controls
- Partnered on PazaBench and supported fine-tuned speech models for African languages
Outcomes & evidence
- ML-ready speech data and quality-control processes for languages with limited suitable vendor datasets
- Support for ASR benchmarking and fine-tuned models for Swahili, Dholuo, Kalenjin, Kikuyu, Maasai, and Somali
Approach & methods
- Started with dataset-gap assessment, then coordinated speaker recruitment, collection, labeling, and annotation.
- Created collection protocols and quality checks where suitable vendor datasets did not exist.
- Supported benchmarking across languages and the review of fine-tuned models, considering accuracy, data requirements, and inference efficiency.
Scope & limitations
- This page describes the research scope and contribution, not unpublished dataset sizes, comparative model scores, or internal experimental findings.
- No language-specific performance claim is presented as a general guarantee.