---
title: "How we source the voice data we train on - Helix"
description: "Every dataset Helix uses to train and evaluate its voice models: where it came from, what it is used for, our legal basis, and the limits of what we can do."
source: "/privacy/training-data"
updated: "2026-09-17"
---

# How we source the voice data we train on

Last updated 17 September 2026 - Version 1.0

Helix builds voice-based age assurance. To do that we have to train machine-learning models on recordings of real human voices, including children's voices. That is not incidental to what we do, it is the core of it, and we think anyone whose voice might be in one of these datasets is entitled to know how we got it and what we do with it.

This statement describes every dataset we use to train and evaluate our models that contains recordings of people we cannot contact. It is published because we cannot tell those people individually, and we would rather be transparent at the level we can manage than say nothing at all. Recordings we collect ourselves, from people who deal with us directly, are described at the end for completeness.

## Why we need voice data at all

We train three models.

An **age-estimation model**, which listens to a short recording and estimates which of five age bands the speaker falls into. A **synthetic-voice detection model**, which works out whether a recording is a real human speaking or something generated by software. And a **replay-detection model**, which works out whether a recording is being played back to a microphone rather than spoken live.

A model that has only ever heard adults cannot tell an adult from a child. It would perform worst at exactly the ages where getting it wrong matters most, and we would have no way of showing a certification body that it works fairly across different groups of speakers. So children's speech data is a requirement of building this responsibly, not a convenience.

## The datasets we use

### Licensed children's speech corpus

**Source.** Nexdata, a specialist speech-data provider.

**What it contains.** Recordings from approximately 219 speakers, totalling around 9.2 hours of audio, with each recording labelled with the speaker's age.

**How it was collected.** By Nexdata, not by us. Nexdata has warranted to us in writing that it obtained consent from participants, or their parents or guardians, at the point of collection, for use in speech-technology development and for onward licensing to companies like ours. We hold that written warranty and we relied on it in deciding whether our use is fair.

**What we use it for.** Training and evaluating the age-estimation model. Nothing else.

**What we do not do with it.** We do not use it to identify anyone. We do not build a voiceprint, speaker profile or speaker gallery from it. We do not attempt to match a recording in it against any other recording, in this dataset or anywhere else. We do not disclose it to our customers or to anyone else, we do not sell or license it out, and we do not contribute it to any dataset exchange or cooperative.

### Synthetic speech corpora

**Source.** Generated using text-to-speech and voice-conversion tools.

**What they contain.** Artificially produced audio. These are not recordings of real people and contain no personal information.

**What we use them for.** Evaluating the synthetic-voice detection model, so that it can learn to recognise generated audio. Synthetic speech supplements real speech in evaluation; it does not replace it for the age model, because a model tuned to synthetic artefacts does not generalise to real voices.

### Presentation-attack corpora

**Source.** Research corpora used across the speech-security field for testing whether a system can be fooled by recorded or synthetic speech. We are completing our documentation of each corpus we use and will name them, with their versions and licence terms, in the next revision of this statement.

**What they contain.** Recordings of adult speakers, together with recorded playbacks and synthetic imitations of those speakers, labelled as genuine or attack. The original speakers were recorded by the research groups that built the corpora, under those groups' own participant terms, and we have no relationship with them.

**What we use them for.** Evaluating the synthetic-voice and replay-detection models. They are not used to train or evaluate the age-estimation model.

**What we do not do with them.** The same as for the children's corpus: no identification, no voiceprints, no matching, no onward disclosure.

### Recordings collected directly by Helix

Where we collect recordings ourselves, we do so from paid adult participants who are told what the recording is for before they take part, who give consent directly to Helix, and who can withdraw it. Because these participants have a direct relationship with us, they receive full notice at the point of collection and can exercise their rights against us in the ordinary way. This collection is adults only by design. At the time of publication it consists of a small initial batch recorded while testing our collection process.

## Our legal basis

For the licensed and research datasets, our lawful basis under UK GDPR is **legitimate interests, Article 6(1)(f)**.

We want to be precise about the role consent plays, because it is easy to get this wrong. The people in the Nexdata corpus gave their consent to Nexdata. Consent has to be given to the organisation relying on it, and has to be withdrawable against that organisation. These speakers have no relationship with us and no way to withdraw anything against us, so their consent cannot be our legal basis and we do not claim it is.

What that consent does do is tell us that this kind of onward use was within the scope of what people were told when they were recorded. That matters to the fairness assessment we have to carry out, and it is one of the reasons we concluded that our interest in building this does not override the interests of the people whose voices are in the dataset. Our full assessment, including the parts that weigh against us, is documented internally and is available to our auditors and our certification body.

We concluded that this processing is not special-category data under Article 9(1) UK GDPR, because we do not process these recordings for the purpose of identifying anyone. That conclusion turns on purpose, and the purpose here is fitting model parameters against age labels, or against genuine-or-attack labels. If that ever changed, so would the legal position, and we treat any such change as requiring a fresh assessment before it happens.

## Why we do not contact people individually

Article 14 UK GDPR normally requires an organisation that obtains personal data from somewhere other than the individual to tell that individual about it. We do not do that here, and we rely on the exemption at Article 14(5)(b) for processing where individual notice would involve disproportionate effort.

The reason is specific rather than convenient. We hold no names, no contact details and no identifiers for the speakers in these corpora, and we have deliberately not sought any. To contact them individually we would first have to acquire identifying information about people we currently cannot identify, which would increase the intrusion rather than reduce it.

Publishing this statement is the measure we take instead. It is not a perfect substitute for individual notice and we do not present it as one.

## How long we keep it, and how it is protected

Training datasets are kept for as long as they are needed to develop and evaluate our models, subject to a written retention policy tied to that purpose and a documented review at least every 24 months. We do not keep children's voice data indefinitely, and the US Children's Online Privacy Protection Act prohibits doing so.

The data is held in an encrypted store with restricted access. Training and evaluation splits are locked and the lineage recorded, both because a certification scheme requires it and because it prevents a model being tested on data it was trained on.

## Your rights, and the limits of what we can do

If you believe a recording of you or your child is in a dataset we license, write to us at <privacy@helix.id> and we will work with the provider to trace and remove it.

Two honest limitations.

**We cannot search for you.** We hold no identifiers for these speakers, so we cannot look up "your" recordings the way we could look up an account. Tracing a recording means going back to the provider that collected it.

**Deleting a recording does not untrain a model.** If a recording is removed from our dataset, it stops being used in future training. It does not reverse the influence it has already had on a model that was trained on it. Only retraining does that. We would rather tell you this than let you assume a deletion request achieves more than it does.

You can also object to our processing under Article 21 UK GDPR. We recognise that in practice this right is difficult to exercise against us for exactly the reason above, and we have recorded that as a genuine weakness rather than reasoning it away.

## Changes to this statement

We review this statement at least every 24 months, and whenever we introduce a new dataset or change what an existing one is used for. Changes are approved before publication and previous versions are retained and available on request.

## Contact

<privacy@helix.id>

H3lix AI Ltd. 1301 N Broadway St, #32355 Los Angeles, CA 90012 United States

Our UK and EU representative under Article 27 is Prighter. EU: Prighter EU Rep GmbH, Schellinggasse 3/10, 1010 Vienna, Austria. UK: Prighter Ltd, 20 Mortlake High Street, London, SW14 8JN, United Kingdom. Requests through <https://app.prighter.com/portal/helix>, quoting reference ID-19548739993.

Our Privacy Notice at <https://helix.id/privacy> explains everything else about how we handle personal information.


## Sitemap

See the full [sitemap](https://helix.id/sitemap.md) for all pages.
