Professional Documents
Culture Documents
Abstract
1. Principles of speaker recognition
2. Speech feature extraction
3. Mfcc processor
4. Feature matching
5. Clustering the training vectors
6. Conclusion
Bibliography
ABSTRACT
Speaker recognition is the process of automatically recognizing who is
speaking on the basis of individual information included in speech waves. This
technique makes it possible to user's voice to verify their identity and control
access to services such as voice dialing, banking by telephone, telephone
shopping, database access services, information services, voice mail, security
control for confidential information areas, and remote access to computers.
S im ilarity
R e fere nc e
m odel M ax im u m Id en tificatio n
Inp u t F e atu re
(S p ea ke r # 1 ) se lectio n re su lt
sp eec h ex trac tio n
(S p e ake r ID )
S im ilarity
R efere nc e
m o d el
(S p ea ke r # N )
V e rifica tio n
In p u t F e a tu re resu lt
S im ilarity D e c isio n
sp ee c h e x tra ctio n (A c ce p t/R e je c t)
R e fere n ce
S p e a k e r ID T h resh o ld
m odel
(# M ) (S p e ak e r # M )
2. Speech Feature Extraction
The purpose of this module is to convert the speech waveform to some
type of parametric representation (at a considerably lower information rate) for
further analysis and processing. This is often referred as the signal-processing
front end.
The speech signal is a slowly timed varying signal (it is called quasi-
stationary). An example of speech signal is shown in Figure. When
examined over a sufficiently short period of time (between 5 and 100 msec),
its characteristics are fairly stationary. However, over long periods of time (on
the order of 1/5 seconds or more) the signal characteristic change to reflect the
different speech sounds being spoken. Therefore, short-time spectral analysis
is the most common way to characterize the speech signal.
0.3
0.25
0.2
0.15
0.1
0.05
-0 . 0 5
-0 . 1
-0 . 1 5
-0 . 2
0 5 10 15 20 25 30 35 40
m s ec
S peaker 1
c entroid
s am ple VQ dis tortion
S peaker 2
c entroid
s am ple
while the triangles are from the speaker 2. In the training phase, a
speaker-specific VQ codebook is generated for each known speaker by
clustering his/her training acoustic vectors. The result codewords (centroids)
are shown in Figure by black circles and black triangles for speaker 1 and 2,
respectively. The distance from a vector to the closest codeword of a
codebook is called a VQ-distortion. In the recognition phase, an input
utterance of an unknown voice is “vector-quantized” using each trained
codebook and the total VQ distortion is computed. The speaker corresponding
to the VQ codebook with smallest total distortion is identified.
One speaker can be discriminated from another based of the location of
centroids
5. Clustering the Training Vectors
After the enrolment session, the acoustic vectors extracted from input
speech of a speaker provide a set of training vectors. As described above, the
next important step is to build a speaker-specific VQ codebook for this
speaker using those training vectors. There is a well-know algorithm, namely
LBG algorithm [Linde, Buzo and Gray, 1980], for clustering a set of L
training vectors into a set of M codebook vectors. The algorithm is formally
implemented by the following recursive procedure:
Design a 1-vector codebook; this is the centroid of the entire set of
training vectors (hence, no iteration is required here).
Double the size of the codebook by splitting each current codebook yn
according to the rule
Y+n = Yn (1+ ), -
Yn = Yn (1-)
Find
centroid
Yes No
m<M Stop
Split each
centroid
m = 2*m
Cluster
vectors
Find
centroids
Compute D
(distortion)
No D '−D Yes
D’ = D <ε
D
Bibliography:
[1] L.R. Rabiner and B.H. Juang, Fundamentals of Speech Recognition,
Prentice-Hall, Englewood Cliffs, N.J., 1993.
[2] L.R Rabiner and R.W. Schafer, Digital Processing of Speech Signals,
Prentice-Hall, Englewood Cliffs, N.J., 1978.