Voice biometrics
Each word that someone speaks is composed of individual sounds called phonemes. The language, culture, and geographical location of each individual speaker determines how many unique sounds that speaker uses when talking. In English, speakers typically use 40 different phonemes when talking.
The actual sound of each phoneme is determined by a speaker’s physical characteristics (vocal chords, nasal passages, and so on.) as well as behavioral characteristics (accent, enunciation, and so on). For an individual speaker, a sentence of speech containing a variety of phonemes spoken with tone, strength and duration, is unique to that speaker. Voice biometrics uses these unique characteristics to establish a voiceprint identity.
This is the basis of voice biometrics. Even though a speaker says words with slight variations, each word generally has the same characteristics. The biometrics system compares spoken utterances with previously enrolled voiceprints, and determines whether they are spoken by the same speaker.
The confidence of a verification match depends partly on how many unique phonemes are captured in recorded utterances. Generally, longer recordings contain more phonemes, and improve the security of the system and the strength of the verification.
The system expects variance in each speaker’s voice. For example, a speaker with a sore throat might change the pitch and speed of their words. Regardless, no speaker can change all characteristics of their voice, and although the system might return lower confidence scores for voiceprint matches, it is still likely that true users are accepted and false users are rejected.
Basic voice biometrics operations
The following table describes the basic operations of the voice biometrics system:
| Operation | Description |
|---|---|
| Verification | Gatekeeper confirms the identity of someone by comparing the caller’s voice to a previously enrolled voiceprint. Text-independent verification uses conversational audio of the speaker’s voice regardless of the content spoken. Text-dependent verification uses a passphrase to identify the speaker. Typically, the passphrase is the same for all speakers (known as a common or shared passphrase). |
| Identification | Gatekeeper discovers the identity of someone by comparing the caller’s voice to a group of voiceprints. |
| Fraud detection | Gatekeeper detects a fraud attempt by comparing the caller’s voice to a watchlist of fraudster voiceprints. |