Voice Playbook The Tech Innovators Network Voice Playbook is a practical guide for building inclusive voice technologies that serve underrepresented languages. It provides a structured framework for collecting, managing, and sharing voice data through community-driven, ethical, and open approaches.Developed with insights from projects such as the Africa Next Voices initiative, the Playbook empowers communities, technologists, and policymakers to collaborate in preserving languages, reducing digital inequality, and shaping the future of speech technology.It is intended as both a reference manual and a living resource, evolving alongside the needs of the communities it supports. Introduction 1.1 The Rise of Voice Technologies Voice has become one of the most powerful interfaces between humans and technology. Around the world, people increasingly interact with digital systems through speech—asking questions, dictating messages, navigating services, and accessing information without needing to type or read. Voice-enabled technologies now support a wide range of applications: virtual assistants such as Siri automated call centers and customer support systems real-time transcription and subtitling speech-based translation tools voice interfaces for mobile applications accessibility tools for people with disabilities Behind each of these systems lies a critical component: large speech datasets . These datasets allow machine learning systems to learn how languages sound in real-world environments. They capture pronunciation patterns, accents, dialectal differences, and the natural rhythm of speech. Over the past decade, advances in machine learning have dramatically improved speech recognition and synthesis systems. However, these improvements have largely benefited high-resource languages such as English, Mandarin, Spanish, and French. Many of the world’s languages remain absent from these technological advances. 1.2 The Global Speech Data Divide Modern speech recognition systems require enormous amounts of training data. State-of-the-art systems often rely on thousands of hours of annotated speech recordings in order to achieve high accuracy. Unfortunately, most languages do not have such resources. Of the roughly 7,000 languages spoken globally , only a small number have the large, well-curated speech datasets necessary to build robust speech technologies. This disparity creates what researchers often refer to as the speech data divide . Languages with abundant data continue to benefit from improved AI tools, while languages with little or no digital data remain excluded from technological innovation. For speakers of these underrepresented languages, the consequences are significant: voice assistants cannot understand them automated services are unavailable in their language speech-to-text systems perform poorly or not at all digital platforms fail to recognize their linguistic identities As voice interfaces become a primary way people interact with technology, this divide risks deepening global digital inequalities . 1.3 African Languages and the Data Gap Africa is the most linguistically diverse continent in the world. It is estimated that over 2,000 languages are spoken across the continent. These languages belong to several major language families. Many African languages are widely spoken by millions of people, yet remain extremely underrepresented in digital datasets . Even widely spoken languages such as Yoruba, Hausa, Amharic, and Swahili often have far fewer digital speech resources compared to European or Asian languages. The situation is even more challenging for languages with fewer speakers. Languages spoken by less than one million people often have little to no available speech data, making it nearly impossible to build speech recognition systems for them. Several factors contribute to this gap: Limited funding for language resource development Scarcity of technical infrastructure for data collection Lack of standardized methodologies for collecting speech datasets Low participation of local communities in AI development processes Barriers to publishing open datasets As a result, many African languages remain digitally invisible . 1.4 Why Speech Data Matters Speech datasets form the foundation for multiple technologies that can benefit communities across Africa. These technologies include: a) Automatic Speech Recognition (ASR) ASR systems convert spoken language into written text. These systems enable: voice search transcription services automated customer service accessibility tools b) Text-to-Speech (TTS) TTS systems convert text into natural-sounding speech. These technologies support: audiobooks assistive technologies voice assistants language learning applications c) Speech Translation Speech translation systems allow spoken communication to be translated into other languages in real time. d) Conversational AI Speech datasets also power chatbots and conversational agents that interact with users through voice. When speech technologies support local languages, they can improve access to services in key areas such as: healthcare agriculture education disaster response financial inclusion public information services However, without speech datasets in these languages, such technologies cannot be developed. 1.5 Community-Driven Speech Data Collection In recent years, a growing number of initiatives have begun addressing the speech data gap for African languages. One of the most promising approaches is c ommunity-driven data collection . Rather than relying solely on centralized research institutions, community-driven approaches engage native speakers, local organizations, and grassroots networks to collect speech recordings. Projects such as AfriSpeech-200 h ave demonstrated the power of community participation in building large-scale datasets. AfriSpeech collected speech recordings from speakers across multiple African countries to create a pan-African accented speech dataset used for speech recognition research. Similarly, the African Voices Kenya project gathered approximately 3 0 00 hours of speech data across five Kenya languages through community engagement and ethical data collection practices. These initiatives show that d istributed community participation can generate large and diverse speech datasets , even in resource-constrained environments. Community-based approaches offer several advantages: they capture authentic speech patterns they include diverse dialects and accents they empower speakers to contribute to digital representation of their languages they create shared ownership of language resources Overview Purpose of the Voice Playbook This playbook provides practical, ethical, and scalable guidance for collecting voice datasets for African languages , particularly low-resource languages with limited digital presence . The playbook enables: Community-led voice dataset creation Standardized speech dataset methodologies Ethical and consent-based recording Scalable workflows for NLP training Reusable processes for future language initiatives The playbook supports Automatic Speech Recognition (ASR) , Text-to-Speech (TTS) , and multimodal language models . African languages remain severely underrepresented in speech datasets, limiting their participation in modern AI systems. Many African languages lack large-scale audio datasets necessary for training speech technologies. This playbook lowers barriers by enabling grassroots communities, universities, NGOs, and language activists to collect high-quality speech data. MODULE 1 How to Use this Playbook This playbook is for anyone who wants to collect voice data in an African language, whether the project is: Small — 50–100 speakers Medium — 500–1,000 speakers Large — thousands of contributors Academic Community-led Government-led NGO-led Commercial Open-source Designed for ASR/TTS Intended for language documentation Intended for AI research It is designed to answer a simple question: “If I want to collect voice data in my language, what exactly do I need to do?” The playbook takes you through the entire lifecycle: Plan → Design → Prepare → Mobilize → Recruit → Consent → Train → Record → Validate → Transcribe → QA → Pay → Document → Release → Sustain Voice collection is not simply a recording exercise. It is an operational system involving people, language expertise, community relationships, technology, quality control, ethics and data management. The AfriVoices-KE project organized these functions through language leads, resource persons, mobilizers, contributors, validators, transcribers, super-reviewers and administrators. PART I — BEFORE YOU COLLECT ANYTHING 1. Define Why You Are Collecting Voice Data Before recruiting anyone, answer: 1.1 What will the data be used for? Examples: Automatic Speech Recognition Speech-to-text Text-to-Speech Speech translation Voice assistants Language documentation Linguistic research Benchmarking Conversational AI Digital government Healthcare Agriculture Education Your intended use determines what type of recordings you need. 1.2 What kind of speech do you need? There are two basic approaches used in the reference project. a) Scripted speech A participant reads a prepared sentence. b) Unscripted speech A participant responds naturally to a question or prompt. The reference project deliberately combined the two, targeting approximately 75% unscripted and 25% scripted speech. Practical recommendation If your goal is a general-purpose speech dataset, consider collecting both. Type What it gives you Scripted Controlled sentences, predictable content Unscripted Natural speech, dialects, spontaneous phrasing Both Broader coverage Do not assume that one approach automatically replaces the other. 2. Decide How Much Data You Need Start with the target rather than starting with recruitment. Define: a)Target number of speakers Example: 1,000 speakers b)Target hours Example : 500 hours c)Target recordings per person Example:  30 minutes validated speech per contributor d)Target speech type Example:  75% unscripted / 25% scripted e)Target geographic coverage Example : 10 counties f) Target demographic representation Example:  Balanced across gender and five age groups. The Africa Next Voices project targeted approximately 3,000 hours across five languages, with different targets per language. g) Create a simple target sheet Metric Target Speakers 1,000 Total hours 500 Scripted 125 hrs Unscripted 375 hrs Regions 10 Dialects 4 Female 50% Male 50% Age groups 5 Your numbers will depend on your project. The table is a planning tool, not a universal standard . 3. Define Your Language Do not simply write: “We are collecting Language X.” Ask: Which variety? Which dialects? Which geographic areas? Which communities? Is the language internally diverse? Are some varieties more widely used? Are some varieties endangered? Which varieties are appropriate for the intended use? Example : The AFN project explicitly considered dialect variation in languages including Kalenjin, Maasai, Somali, Dholuo and Kikuyu. Create a language profile Record: Language: Language family: Country/countries: Main regions: Dialect groups: Estimated speaker distribution: Existing digital resources: Existing speech datasets: Native-speaker experts: Relevant institutions: 4. Define Your Dialect Strategy This is one of the most important steps. A dataset can contain thousands of hours and still be poorly representative if most speakers come from one location or dialect. Create a dialect matrix: Dialect Geography Target speakers Target hours Dialect A Region 1 200 100 Dialect B Region 2 200 100 Dialect C Region 3 200 100 Dialect D Region 4 200 100 The Africa Next Voices project used explicit dialect targets and geographic recruitment rather than treating each language as homogeneous. PART II — DESIGN THE DATASET 5. Choose Scripted vs Unscripted Collection 5.1 Scripted Collection Participants read prepared sentences. a) Use scripted data when you need: Consistent text/audio pairs Controlled vocabulary Sentence-level alignment Clear transcription Specific words or phrases Named entities Domain-specific terminology b) Advantages Fast to collect Easy to validate Predictable content Easy to align with text c) Limitations The reference report notes that scripted speech can omit natural speech phenomena such as: Fillers False starts Slang Code-switching Natural prosody Informal speech It can also introduce bias toward formal language. 6. Unscripted Collection Unscripted collection asks contributors to speak naturally in response to a prompt. For example: “Describe what you normally do when you wake up in the morning.” or: “Tell us about a memorable journey you have taken.” or: Show an image and ask the participant to describe what is happening. The reference project used both text and image prompts to elicit spontaneous speech. Advantages Captures: Natural phrasing Accent Dialect Spontaneous speech Informal vocabulary Prosody Natural discourse Challenges Unscripted recordings require substantially more: Validation Transcription QA Prompt design Human review The reference project therefore used a multi-level QA and transcription process for this data. 7. Decide Your Domains Do not collect speech around one topic only. The reference project used multiple domains, including: Agriculture and Food Everyday Scenarios Financial Transactions Digital Government Services Named Entity Recognition Role Play Extempore Stories Healthcare News and Media Education and Technology Customer Care Create your own domain matrix Domain Scripted Unscripted Target hours Everyday life ✓ ✓ 50 Agriculture ✓ ✓ 100 Healthcare ✓ ✓ 50 Education ✓ ✓ 50 Government ✓ ✓ 50 Stories   ✓ 50 Customer care ✓ ✓ 50 Other ✓ ✓ 100 The purpose is to expose the eventual model to varied vocabulary and contexts. 8. Develop Your Scripted Sentences Your scripted sentences should not simply be random translations. Build them from: Existing corpora Existing language resources Domain-specific material Locally generated sentences Carefully reviewed translations The reference project combined existing corpora with translated and newly generated material. Every sentence should be checked for: Naturalness Correct grammar Correct spelling Correct terminology Cultural appropriateness Dialect Names Places Numbers Dates Abbreviations Important rule Do not assume an English sentence can simply be translated word-for-word. The reference project specifically encountered problems with polysemous words, inconsistent translations and concepts that lacked direct equivalents in some languages. 9. Develop Unscripted Prompts Good prompts produce speech. Bad prompt: “Do you like farming?” Possible response: “Yes.” Good prompt: “Tell us about the farming activities that people in your community usually do during the rainy season.” This encourages a longer response. Use several prompt types a)Text prompts A written question. b) Image prompts An image plus a question. c) Story prompts Ask participants to tell a story. d) Scenario prompts Ask participants to explain what they would do. e) Role-play prompts Example: “Imagine you are calling a customer-care agent because your mobile money transaction failed.”