基于Gammatone滤波器组的听觉特征提取

doi:10.3969/j.issn.1000-3428.2012.21.045

计算机工程 ›› 2012, Vol. 38 ›› Issue (21): 168-170,174.

基于Gammatone滤波器组的听觉特征提取

胡峰松^1,2，曹孝玉¹

(1. 湖南大学信息科学与工程学院，长沙 410082；2. 北京师范大学管理学院，北京 100875)

收稿日期:2012-02-14 出版日期:2012-11-05 发布日期:2012-11-02
作者简介:胡峰松(1969－)，男，副教授、博士，主研方向：语音识别，人脸识别；曹孝玉，硕士研究生

Auditory Feature Extraction Based on Gammatone Filter Bank

HU Feng-song ^1,2, CAO Xiao-yu ¹

(1. College of Information Science and Engineering, Hunan University, Changsha 410082, China; 2. School of Management, Beijing Normal University, Beijing 100875, China)

Received:2012-02-14 Online:2012-11-05 Published:2012-11-02

摘要/Abstract

摘要： 目前主流说话人特征参数在噪声环境中的鲁棒性较差。为此，提出一种可用于说话人识别的听觉倒谱特征系数。分析人耳听觉模型的工作机理，采用Gammatone滤波器组代替传统的三角滤波器组模拟人耳耳蜗的听觉模型，用指数压缩代替固定的对数压缩，模拟人耳听觉模型处理信号的非线性特性。在基于高斯混合模型分类器的识别算法下进行仿真实验，结果表明，该听觉特征具有比梅尔频率倒谱系数和线性预测倒谱系数更好的抗噪声能力。

关键词: 说话人识别, 特征提取, Gammatone 滤波器, 听觉模型, 倒谱系数, 鲁棒性

Abstract: Aiming at the problem that speaker’s feature coefficients have poor robustness in noise environment, this paper proposes an auditory cepstral coefficient for speaker recognition. It analyzes the working mechanism of the human auditory model, simulates the auditory model of human ear cochlea by Gammatone filter banks replaces the traditional triangular filter banks. Based on the nonlinear signal processing capability of human auditory model, exponential compression is used instead of the fixed logarithm compression. Simulation experiment is conducted based on Gaussian Mixed Model(GMM) recognition algorithm. Experimental results show that the auditory feature has better noise robustness than Mel Frequency Cepstral Coefficient(MFCC) and Linear Prediction Cepstral Coefficient(LPCC).

Key words: speaker recognition, feature extraction, Gammatone filter, auditory model, cepstral coefficient, robustness

中图分类号:

TP391

胡峰松, 曹孝玉. 基于Gammatone滤波器组的听觉特征提取[J]. 计算机工程, 2012, 38(21): 168-170,174.

HU Feng-Song, CAO Xiao-Yu. Auditory Feature Extraction Based on Gammatone Filter Bank[J]. Computer Engineering, 2012, 38(21): 168-170,174.

https://www.ecice06.com/CN/Y2012/V38/I21/168

[1]	沈忱, 何勇, 彭安浪. 鲁棒物联网多维时序数据预测方法[J]. 计算机工程, 2025, 51(4): 107-118.
[2]	董红亮, 钮焱, 孙杨, 李军. 基于记忆胶囊与注意力的语音情感识别[J]. 计算机工程, 2025, 51(4): 169-177.
[3]	孙义康, 高建华. 基于卷积神经网络和长短期记忆的死代码检测方法[J]. 计算机工程, 2025, 51(2): 223-237.
[4]	赵宏, 宋馥荣, 李文改. 基于SE-AdvGAN的图像对抗样本生成方法研究[J]. 计算机工程, 2025, 51(2): 300-311.
[5]	许明, 屈泰澎, 姜彦吉. 改进YOLOv7在复杂场景下的交通标志检测算法[J]. 计算机工程, 2025, 51(2): 335-343.
[6]	张新波, 张雪英, 黄丽霞, 陈桂军. 基于半监督深度自编码网络的分类算法及应用[J]. 计算机工程, 2025, 51(1): 71-80.
[7]	喻勇涛, 孙奥, 李昂, 朱琳琳. 基于孪生网络的分类器输出重复性优化方法[J]. 计算机工程, 2025, 51(1): 118-127.
[8]	郑秋梅, 赵丹, 牛薇薇, 林超. 基于多通道的彩色图像多重水印算法[J]. 计算机工程, 2024, 50(9): 246-254.
[9]	李维刚, 厉许昌, 田志强, 李金灵. 基于自蒸馏框架的点云分类及其鲁棒性研究[J]. 计算机工程, 2024, 50(9): 72-81.
[10]	赵俊涛, 李陶深, 卢志翔. 基于最优近邻的局部保持投影方法[J]. 计算机工程, 2024, 50(9): 161-168.
[11]	钱清, 龙永, 蒋忠远, 段春红, 王宏. 基于深度强化学习的自适应图像隐写算法[J]. 计算机工程, 2024, 50(8): 319-327.
[12]	胡庆. 多尺度融合与双输出U-Net网络的行人重识别[J]. 计算机工程, 2024, 50(6): 102-109.
[13]	顾永跟, 高凌轩, 吴小红, 陶杰. 非独立同分布下联邦半监督学习的数据分享研究[J]. 计算机工程, 2024, 50(6): 188-196.
[14]	梁松林, 林伟, 王珏, 杨庆. 面向后渗透攻击行为的网络恶意流量检测研究[J]. 计算机工程, 2024, 50(5): 128-138.
[15]	李振鲁, 黄威, 孙锴. 复杂环境下的轻量化道路目标识别算法研究[J]. 计算机工程, 2024, 50(4): 219-227.

选择文件类型/文献管理软件名称

选择包含的内容

基于Gammatone滤波器组的听觉特征提取

Auditory Feature Extraction Based on Gammatone Filter Bank

PDF

可视化

被引次数

摘要/Abstract

引用本文

使用本文

参考文献

相关文章 15

编辑推荐

Metrics

本文评价

模态框（Modal）标题

选择文件类型/文献管理软件名称

选择包含的内容

基于Gammatone滤波器组的听觉特征提取

Auditory Feature Extraction Based on Gammatone Filter Bank

PDF

可视化

被引次数

摘要/Abstract

引用本文

使用本文

参考文献

相关文章 15

编辑推荐

Metrics

本文评价