# Voice Playbook

<span>The </span>**Tech Innovators Network Voice Playbook**<span> is a practical guide for building inclusive voice technologies that serve underrepresented languages. It provides a structured framework for </span>**collecting, managing, and sharing voice data**<span> through community-driven, ethical, and open approaches.</span>  
  
<span>Developed with insights from projects such as the </span>**Africa Next Voices initiative**, the Playbook empowers communities, technologists, and policymakers to collaborate in preserving languages, reducing digital inequality, and shaping the future of speech technology.  
  
<span>It is intended as both a </span>**reference manual**<span> and a </span>**living resource**, evolving alongside the needs of the communities it supports.

[  ](https://voice-1.gitbook.io/voice-docs/think-voice-playbook/1.-introduction)

# Introduction

### 1.1 The Rise of Voice Technologies

Voice has become one of the most powerful interfaces between humans and technology. Around the world, people increasingly interact with digital systems through speech—asking questions, dictating messages, navigating services, and accessing information without needing to type or read.

Voice-enabled technologies now support a wide range of applications:

- virtual assistants such as Siri
- automated call centers and customer support systems
- real-time transcription and subtitling
- speech-based translation tools
- voice interfaces for mobile applications
- accessibility tools for people with disabilities

<p class="callout info">Behind each of these systems lies a critical component: <span style="color:rgb(132,63,161);">**large speech datasets**.</span></p>

These datasets allow machine learning systems to learn how languages sound in real-world environments. They capture pronunciation patterns, accents, dialectal differences, and the natural rhythm of speech.

Over the past decade, advances in machine learning have dramatically improved speech recognition and synthesis systems. However, these improvements have largely benefited <span style="color:rgb(132,63,161);">**high-resource languages** </span>such as English, Mandarin, Spanish, and French.

Many of the world’s languages remain absent from these technological advances.

### 1.2 The Global Speech Data Divide

Modern speech recognition systems require enormous amounts of training data. State-of-the-art systems often rely on<span style="color:rgb(132,63,161);"> **thousands of hours of annotated speech recordings**</span> in order to achieve high accuracy.

Unfortunately, most languages do not have such resources.

Of the roughly <span style="color:rgb(132,63,161);">**7,000 languages spoken globally**,</span> only a small number have the large, well-curated speech datasets necessary to build robust speech technologies.

This disparity creates what researchers often refer to as the <span style="color:rgb(132,63,161);">**speech data divide**</span>.

Languages with abundant data continue to benefit from improved AI tools, while languages with little or no digital data remain excluded from technological innovation.

For speakers of these underrepresented languages, the consequences are significant:

- voice assistants cannot understand them
- automated services are unavailable in their language
- speech-to-text systems perform poorly or not at all
- digital platforms fail to recognize their linguistic identities

As voice interfaces become a primary way people interact with technology, this divide risks **deepening global digital inequalities**.

### 1.3 African Languages and the Data Gap

Africa is the most linguistically diverse continent in the world.

It is estimated that <span style="color:rgb(132,63,161);">**over 2,000 languages** </span>are spoken across the continent. These languages belong to several major language families. Many African languages are widely spoken by millions of people, yet remain <span style="color:rgb(132,63,161);">**extremely underrepresented in digital datasets**.</span>

Even widely spoken languages such as Yoruba, Hausa, Amharic, and Swahili often have far fewer digital speech resources compared to European or Asian languages.

The situation is even more challenging for languages with fewer speakers. Languages spoken by<span style="color:rgb(132,63,161);"> **less than one million people**</span> often have little to no available speech data, making it nearly impossible to build speech recognition systems for them.

Several factors contribute to this gap:

1. <span style="color:rgb(132,63,161);">**Limited funding for language resource development**</span>
2. <span style="color:rgb(132,63,161);">**Scarcity of technical infrastructure for data collection**</span>
3. <span style="color:rgb(132,63,161);">**Lack of standardized methodologies for collecting speech datasets**</span>
4. <span style="color:rgb(132,63,161);">**Low participation of local communities in AI development processes**</span>
5. <span style="color:rgb(132,63,161);">**Barriers to publishing open datasets**</span>

As a result, many African languages remain <span style="color:rgb(132,63,161);">**digitally invisible**</span>.

### 1.4 Why Speech Data Matters

Speech datasets form the foundation for multiple technologies that can benefit communities across Africa.

These technologies include:

##### a) Automatic Speech Recognition (ASR)

ASR systems convert spoken language into written text. These systems enable:

- voice search
- transcription services
- automated customer service
- accessibility tools

##### b) Text-to-Speech (TTS)

TTS systems convert text into natural-sounding speech. These technologies support:

- audiobooks
- assistive technologies
- voice assistants
- language learning applications

##### c) Speech Translation

Speech translation systems allow spoken communication to be translated into other languages in real time.

##### d) Conversational AI

Speech datasets also power chatbots and conversational agents that interact with users through voice.

When speech technologies support local languages, they can improve access to services in key areas such as:

- healthcare
- agriculture
- education
- disaster response
- financial inclusion
- public information services

However, without speech datasets in these languages, such technologies cannot be developed.

### 1.5 Community-Driven Speech Data Collection

In recent years, a growing number of initiatives have begun addressing the speech data gap for African languages.

One of the most promising approaches is **c<span style="color:rgb(132,63,161);">ommunity-driven data collection</span>**<span style="color:rgb(132,63,161);">.</span>

Rather than relying solely on centralized research institutions, community-driven approaches engage <span style="color:rgb(132,63,161);">**native speakers, local organizations, and grassroots networks**</span> to collect speech recordings.

Projects such as<span style="color:rgb(132,63,161);"> <span style="background-color:rgb(132,63,161);">**<span style="background-color:rgb(255,255,255);">AfriSpeech-200 </span>**</span></span><span style="background-color:rgb(255,255,255);">h</span>ave demonstrated the power of community participation in building large-scale datasets. AfriSpeech collected speech recordings from speakers across multiple African countries to create a pan-African accented speech dataset used for speech recognition research.

Similarly, the <span style="color:rgb(132,63,161);">**African Voices Kenya** </span>project gathered approximately 3<span style="color:rgb(132,63,161);">0**00 hours of speech data across five Kenya languages**</span> through community engagement and ethical data collection practices.

These initiatives show that **d<span style="color:rgb(132,63,161);">istributed community participation can generate large and diverse speech datasets</span>**, even in resource-constrained environments.

Community-based approaches offer several advantages:

- they capture authentic speech patterns
- they include diverse dialects and accents
- they empower speakers to contribute to digital representation of their languages
- they create shared ownership of language resources

# Overview

### **Purpose of the Voice Playbook**

This playbook provides **practical, ethical, and scalable guidance** for collecting **voice datasets for African languages**, particularly **low-resource languages with limited digital presence**.

The playbook enables:

- Community-led voice dataset creation
- Standardized speech dataset methodologies
- Ethical and consent-based recording
- Scalable workflows for NLP training
- Reusable processes for future language initiatives

The playbook supports **Automatic Speech Recognition (ASR)**, **Text-to-Speech (TTS)**, and **multimodal language models**.

African languages remain severely underrepresented in speech datasets, limiting their participation in modern AI systems. Many African languages lack large-scale audio datasets necessary for training speech technologies.

> **<span style="color:rgb(132,63,161);">This playbook lowers barriers by enabling grassroots communities, universities, NGOs, and language activists to collect high-quality speech data.</span>**

# MODULE 1

# How to Use this Playbook

This playbook is for anyone who wants to collect voice data in an African language, whether the project is:

- Small — 50–100 speakers
- Medium — 500–1,000 speakers
- Large — thousands of contributors
- Academic
- Community-led
- Government-led
- NGO-led
- Commercial
- Open-source
- Designed for ASR/TTS
- Intended for language documentation
- Intended for AI research

It is designed to answer a simple question:

<span style="background-color: rgb(236, 202, 250);">***“If I want to collect voice data in my language, what exactly do I need to do?”***</span>

The playbook takes you through the entire lifecycle:

> ##### Plan → Design → Prepare → Mobilize → Recruit → Consent → Train → Record → Validate → Transcribe → QA → Pay → Document → Release → Sustain

  
  
Voice collection is not simply a recording exercise. It is an operational system involving people, language expertise, community relationships, technology, quality control, ethics and data management. The AfriVoices-KE project organized these functions through language leads, resource persons, mobilizers, contributors, validators, transcribers, super-reviewers and administrators.

# PART I — BEFORE YOU COLLECT ANYTHING

### <span style="background-color: rgb(255, 255, 255);">1. Define Why You Are Collecting Voice Data</span>

Before recruiting anyone, answer:

#### 1.1 What will the data be used for?

Examples:

- Automatic Speech Recognition
- Speech-to-text
- Text-to-Speech
- Speech translation
- Voice assistants
- Language documentation
- Linguistic research
- Benchmarking
- Conversational AI
- Digital government
- Healthcare
- Agriculture
- Education

<p class="callout success">Your intended use determines what type of recordings you need.</p>

#### 1.2 What kind of speech do you need?

There are two basic approaches used in the reference project.

**a) Scripted speech**

A participant reads a prepared sentence.

**b) Unscripted speech**

A participant responds naturally to a question or prompt.

The reference project deliberately combined the two, targeting approximately 75% unscripted and 25% scripted speech.

#### Practical recommendation

If your goal is a general-purpose speech dataset, consider collecting both.

<div align="left" dir="ltr" id="bkmrk-type-what-it-gives-y"><table style="width: 61.5476%;"><colgroup><col style="width: 37.9523%;" width="85"></col><col style="width: 62.0477%;" width="313"></col></colgroup><tbody><tr><td><p class="callout info">Type</p>

</td><td><p class="callout info">What it gives you</p>

</td></tr><tr><td><p class="callout info">Scripted</p>

</td><td><p class="callout info">Controlled sentences, predictable content</p>

</td></tr><tr><td><p class="callout info">Unscripted</p>

</td><td><p class="callout info">Natural speech, dialects, spontaneous phrasing</p>

</td></tr><tr><td><p class="callout info">Both</p>

</td><td><p class="callout info">Broader coverage</p>

</td></tr></tbody></table>

</div>Do not assume that one approach automatically replaces the other.

#### 2. Decide How Much Data You Need

Start with the target rather than starting with recruitment.

Define:

##### a)Target number of speakers

<span style="background-color: rgb(236, 202, 250);">Example: </span>1,000 speakers

##### b)Target hours

<span style="background-color: rgb(236, 202, 250);">Example</span>: 500 hours

##### c)Target recordings per person

<span style="background-color: rgb(236, 202, 250);">Example:</span> 30 minutes validated speech per contributor

##### d)Target speech type

<span style="background-color: rgb(236, 202, 250);">Example:</span> 75% unscripted / 25% scripted

##### e)Target geographic coverage

<span style="background-color: rgb(236, 202, 250);">Example</span>: 10 counties

##### f) Target demographic representation

<span style="background-color: rgb(236, 202, 250);">Example:</span> Balanced across gender and five age groups.

> The Africa Next Voices project targeted approximately 3,000 hours across five languages, with different targets per language.

##### g) Create a simple target sheet

<div align="left" dir="ltr" id="bkmrk-metric-target-speake"><table style="width: 41.1905%;"><colgroup><col style="width: 58.7879%;" width="89"></col><col style="width: 41.2121%;" width="63"></col></colgroup><tbody><tr><td><p class="callout success">Metric</p>

</td><td><p class="callout success">Target</p>

</td></tr><tr><td><p class="callout success">Speakers</p>

</td><td><p class="callout success">1,000</p>

</td></tr><tr><td><p class="callout success">Total hours</p>

</td><td><p class="callout success">500</p>

</td></tr><tr><td><p class="callout success">Scripted</p>

</td><td><p class="callout success">125 hrs</p>

</td></tr><tr><td><p class="callout success">Unscripted</p>

</td><td><p class="callout success">375 hrs</p>

</td></tr><tr><td><p class="callout success">Regions</p>

</td><td><p class="callout success">10</p>

</td></tr><tr><td><p class="callout success">Dialects</p>

</td><td><p class="callout success">4</p>

</td></tr><tr><td><p class="callout success">Female</p>

</td><td><p class="callout success">50%</p>

</td></tr><tr><td><p class="callout success">Male</p>

</td><td><p class="callout success">50%</p>

</td></tr><tr><td><p class="callout success">Age groups</p>

</td><td><p class="callout success">5</p>

</td></tr></tbody></table>

</div>**<span style="background-color: rgb(236, 202, 250);">*Your numbers will depend on your project. The table is a planning tool, not a universal standard*</span>**.

#### 3. Define Your Language

Do not simply write:

<p class="callout success">**“We are collecting Language X.”**</p>

Ask:

- Which variety?
- Which dialects?
- Which geographic areas?
- Which communities?
- Is the language internally diverse?
- Are some varieties more widely used?
- Are some varieties endangered?
- Which varieties are appropriate for the intended use?

> Example : The AFN project explicitly considered dialect variation in languages including Kalenjin, Maasai, Somali, Dholuo and Kikuyu.

#### Create a language profile

Record:

1. Language:
2. Language family:
3. Country/countries:
4. Main regions:
5. Dialect groups:
6. Estimated speaker distribution:
7. Existing digital resources:
8. Existing speech datasets:
9. Native-speaker experts:
10. Relevant institutions:

#### 4. Define Your Dialect Strategy

This is one of the most important steps.

A dataset can contain thousands of hours and still be poorly representative if most speakers come from one location or dialect.

Create a dialect matrix:

<div align="left" dir="ltr" id="bkmrk-dialect-geography-ta"><table style="width: 71.6667%;"><colgroup><col style="width: 19.5431%;" width="77"></col><col style="width: 23.0964%;" width="91"></col><col style="width: 31.2183%;" width="123"></col><col style="width: 26.1421%;" width="103"></col></colgroup><tbody><tr><td><p class="callout success">Dialect</p>

</td><td><p class="callout success">Geography</p>

</td><td><p class="callout success">Target speakers</p>

</td><td><p class="callout success">Target hours</p>

</td></tr><tr><td><p class="callout success">Dialect A</p>

</td><td><p class="callout success">Region 1</p>

</td><td><p class="callout success">200</p>

</td><td><p class="callout success">100</p>

</td></tr><tr><td><p class="callout success">Dialect B</p>

</td><td><p class="callout success">Region 2</p>

</td><td><p class="callout success">200</p>

</td><td><p class="callout success">100</p>

</td></tr><tr><td><p class="callout success">Dialect C</p>

</td><td><p class="callout success">Region 3</p>

</td><td><p class="callout success">200</p>

</td><td><p class="callout success">100</p>

</td></tr><tr><td><p class="callout success">Dialect D</p>

</td><td><p class="callout success">Region 4</p>

</td><td><p class="callout success">200</p>

</td><td><p class="callout success">100</p>

</td></tr></tbody></table>

</div>> The Africa Next Voices project used explicit dialect targets and geographic recruitment rather than treating each language as homogeneous.

# PART II — DESIGN THE DATASET

### 5. Choose Scripted vs Unscripted Collection

#### 5.1 Scripted Collection

Participants read prepared sentences.

##### a) Use scripted data when you need:

- Consistent text/audio pairs
- Controlled vocabulary
- Sentence-level alignment
- Clear transcription
- Specific words or phrases
- Named entities
- Domain-specific terminology

##### b) Advantages

- Fast to collect
- Easy to validate
- Predictable content
- Easy to align with text

##### c) Limitations

The reference report notes that scripted speech can omit natural speech phenomena such as:

- Fillers
- False starts
- Slang
- Code-switching
- Natural prosody
- Informal speech

It can also introduce bias toward formal language.

### 6. Unscripted Collection

Unscripted collection asks contributors to speak naturally in response to a prompt.

<span style="background-color: rgb(236, 202, 250);">For example:</span>

<span style="color: rgb(132, 63, 161);">***“Describe what you normally do when you wake up in the morning.”***</span>

or:

<span style="color: rgb(132, 63, 161);">***“Tell us about a memorable journey you have taken.”***</span>

or:

*<span style="color: rgb(132, 63, 161);">**Show an image and ask the participant to describe what is happening.**</span>*

> The reference project used both text and image prompts to elicit spontaneous speech.

#### Advantages

Captures:

- Natural phrasing
- Accent
- Dialect
- Spontaneous speech
- Informal vocabulary
- Prosody
- Natural discourse

#### Challenges

Unscripted recordings require substantially more:

- Validation
- Transcription
- QA
- Prompt design
- Human review

The reference project therefore used a multi-level QA and transcription process for this data.

### 7. Decide Your Domains

Do not collect speech around one topic only.

The reference project used multiple domains, including:

- Agriculture and Food
- Everyday Scenarios
- Financial Transactions
- Digital Government Services
- Named Entity Recognition
- Role Play
- Extempore Stories
- Healthcare
- News and Media
- Education and Technology
- Customer Care

#### Create your own domain matrix

<div align="left" dir="ltr" id="bkmrk-domain-scripted-unsc"><table style="width: 71.0714%;"><colgroup><col style="width: 28.836%;" width="108"></col><col style="width: 19.5767%;" width="73"></col><col style="width: 24.3386%;" width="91"></col><col style="width: 27.5132%;" width="103"></col></colgroup><tbody><tr><td><p class="callout success">Domain</p>

</td><td><p class="callout success">Scripted</p>

</td><td><p class="callout success">Unscripted</p>

</td><td><p class="callout success">Target hours</p>

</td></tr><tr><td><p class="callout success">Everyday life</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">50</p>

</td></tr><tr><td><p class="callout success">Agriculture</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">100</p>

</td></tr><tr><td><p class="callout success">Healthcare</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">50</p>

</td></tr><tr><td><p class="callout success">Education</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">50</p>

</td></tr><tr><td><p class="callout success">Government</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">50</p>

</td></tr><tr><td><p class="callout success">Stories</p>

</td><td> </td><td><p class="callout success">✓</p>

</td><td><p class="callout success">50</p>

</td></tr><tr><td><p class="callout success">Customer care</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">50</p>

</td></tr><tr><td><p class="callout success">Other</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">✓</p>

</td><td><p class="callout success">100</p>

</td></tr></tbody></table>

</div>The purpose is to expose the eventual model to varied vocabulary and contexts.

### 8. Develop Your Scripted Sentences

Your scripted sentences should not simply be random translations.

Build them from:

- Existing corpora
- Existing language resources
- Domain-specific material
- Locally generated sentences
- Carefully reviewed translations

The reference project combined existing corpora with translated and newly generated material.

#### Every sentence should be checked for:

- Naturalness
- Correct grammar
- Correct spelling
- Correct terminology
- Cultural appropriateness
- Dialect
- Names
- Places
- Numbers
- Dates
- Abbreviations

#### Important rule

Do not assume an English sentence can simply be translated word-for-word.

The reference project specifically encountered problems with polysemous words, inconsistent translations and concepts that lacked direct equivalents in some languages.

### 9. Develop Unscripted Prompts

Good prompts produce speech.

<span style="background-color: rgb(236, 202, 250);">**Bad prompt:**</span>

##### “Do you like farming?”

<span style="background-color: rgb(236, 202, 250);">Possible response:</span>

“Yes.”

<span style="background-color: rgb(236, 202, 250);">Good prompt:</span>

“Tell us about the farming activities that people in your community usually do during the rainy season.”

This encourages a longer response.

#### Use several prompt types

#### a)Text prompts

A written question.

##### b) Image prompts

An image plus a question.

##### c) Story prompts

Ask participants to tell a story.

##### d) Scenario prompts

Ask participants to explain what they would do.

##### e) Role-play prompts

Example:

<p class="callout info">**“Imagine you are calling a customer-care agent because your mobile money transaction failed.”**</p>