# PART I — BEFORE YOU COLLECT ANYTHING

### <span style="background-color: rgb(255, 255, 255);">1. Define Why You Are Collecting Voice Data</span>

Before recruiting anyone, answer:

#### 1.1 What will the data be used for?

Examples:

- Automatic Speech Recognition
- Speech-to-text
- Text-to-Speech
- Speech translation
- Voice assistants
- Language documentation
- Linguistic research
- Benchmarking
- Conversational AI
- Digital government
- Healthcare
- Agriculture
- Education

<p class="callout success">Your intended use determines what type of recordings you need.</p>

#### 1.2 What kind of speech do you need?

There are two basic approaches used in the reference project.

**a) Scripted speech**

A participant reads a prepared sentence.

**b) Unscripted speech**

A participant responds naturally to a question or prompt.

The reference project deliberately combined the two, targeting approximately 75% unscripted and 25% scripted speech.

#### Practical recommendation

If your goal is a general-purpose speech dataset, consider collecting both.

<div align="left" dir="ltr" id="bkmrk-type-what-it-gives-y"><table style="width: 61.5476%;"><colgroup><col style="width: 37.9523%;" width="85"></col><col style="width: 62.0477%;" width="313"></col></colgroup><tbody><tr><td><p class="callout info">Type</p>

</td><td><p class="callout info">What it gives you</p>

</td></tr><tr><td><p class="callout info">Scripted</p>

</td><td><p class="callout info">Controlled sentences, predictable content</p>

</td></tr><tr><td><p class="callout info">Unscripted</p>

</td><td><p class="callout info">Natural speech, dialects, spontaneous phrasing</p>

</td></tr><tr><td><p class="callout info">Both</p>

</td><td><p class="callout info">Broader coverage</p>

</td></tr></tbody></table>

</div>Do not assume that one approach automatically replaces the other.

#### 2. Decide How Much Data You Need

Start with the target rather than starting with recruitment.

Define:

##### a)Target number of speakers

<span style="background-color: rgb(236, 202, 250);">Example: </span>1,000 speakers

##### b)Target hours

<span style="background-color: rgb(236, 202, 250);">Example</span>: 500 hours

##### c)Target recordings per person

<span style="background-color: rgb(236, 202, 250);">Example:</span> 30 minutes validated speech per contributor

##### d)Target speech type

<span style="background-color: rgb(236, 202, 250);">Example:</span> 75% unscripted / 25% scripted

##### e)Target geographic coverage

<span style="background-color: rgb(236, 202, 250);">Example</span>: 10 counties

##### f) Target demographic representation

<span style="background-color: rgb(236, 202, 250);">Example:</span> Balanced across gender and five age groups.

> The Africa Next Voices project targeted approximately 3,000 hours across five languages, with different targets per language.

##### g) Create a simple target sheet

<div align="left" dir="ltr" id="bkmrk-metric-target-speake"><table style="width: 41.1905%;"><colgroup><col style="width: 58.7879%;" width="89"></col><col style="width: 41.2121%;" width="63"></col></colgroup><tbody><tr><td><p class="callout success">Metric</p>

</td><td><p class="callout success">Target</p>

</td></tr><tr><td><p class="callout success">Speakers</p>

</td><td><p class="callout success">1,000</p>

</td></tr><tr><td><p class="callout success">Total hours</p>

</td><td><p class="callout success">500</p>

</td></tr><tr><td><p class="callout success">Scripted</p>

</td><td><p class="callout success">125 hrs</p>

</td></tr><tr><td><p class="callout success">Unscripted</p>

</td><td><p class="callout success">375 hrs</p>

</td></tr><tr><td><p class="callout success">Regions</p>

</td><td><p class="callout success">10</p>

</td></tr><tr><td><p class="callout success">Dialects</p>

</td><td><p class="callout success">4</p>

</td></tr><tr><td><p class="callout success">Female</p>

</td><td><p class="callout success">50%</p>

</td></tr><tr><td><p class="callout success">Male</p>

</td><td><p class="callout success">50%</p>

</td></tr><tr><td><p class="callout success">Age groups</p>

</td><td><p class="callout success">5</p>

</td></tr></tbody></table>

</div>**<span style="background-color: rgb(236, 202, 250);">*Your numbers will depend on your project. The table is a planning tool, not a universal standard*</span>**.

#### 3. Define Your Language

Do not simply write:

<p class="callout success">**“We are collecting Language X.”**</p>

Ask:

- Which variety?
- Which dialects?
- Which geographic areas?
- Which communities?
- Is the language internally diverse?
- Are some varieties more widely used?
- Are some varieties endangered?
- Which varieties are appropriate for the intended use?

> Example : The AFN project explicitly considered dialect variation in languages including Kalenjin, Maasai, Somali, Dholuo and Kikuyu.

#### Create a language profile

Record:

1. Language:
2. Language family:
3. Country/countries:
4. Main regions:
5. Dialect groups:
6. Estimated speaker distribution:
7. Existing digital resources:
8. Existing speech datasets:
9. Native-speaker experts:
10. Relevant institutions:

#### 4. Define Your Dialect Strategy

This is one of the most important steps.

A dataset can contain thousands of hours and still be poorly representative if most speakers come from one location or dialect.

Create a dialect matrix:

<div align="left" dir="ltr" id="bkmrk-dialect-geography-ta"><table style="width: 71.6667%;"><colgroup><col style="width: 19.5431%;" width="77"></col><col style="width: 23.0964%;" width="91"></col><col style="width: 31.2183%;" width="123"></col><col style="width: 26.1421%;" width="103"></col></colgroup><tbody><tr><td><p class="callout success">Dialect</p>

</td><td><p class="callout success">Geography</p>

</td><td><p class="callout success">Target speakers</p>

</td><td><p class="callout success">Target hours</p>

</td></tr><tr><td><p class="callout success">Dialect A</p>

</td><td><p class="callout success">Region 1</p>

</td><td><p class="callout success">200</p>

</td><td><p class="callout success">100</p>

</td></tr><tr><td><p class="callout success">Dialect B</p>

</td><td><p class="callout success">Region 2</p>

</td><td><p class="callout success">200</p>

</td><td><p class="callout success">100</p>

</td></tr><tr><td><p class="callout success">Dialect C</p>

</td><td><p class="callout success">Region 3</p>

</td><td><p class="callout success">200</p>

</td><td><p class="callout success">100</p>

</td></tr><tr><td><p class="callout success">Dialect D</p>

</td><td><p class="callout success">Region 4</p>

</td><td><p class="callout success">200</p>

</td><td><p class="callout success">100</p>

</td></tr></tbody></table>

</div>> The Africa Next Voices project used explicit dialect targets and geographic recruitment rather than treating each language as homogeneous.