跳到主要内容

音频质量检测 API

评估上传音频的感知质量,并返回总体质量分数、语音检测结果、背景噪声等级、音量以及削波/失真指标。

基础 URL​

https://api.genderrecognition.com

端点​

POST /v1/audio/quality/api

请求标头​

apiKey: YOUR_API_KEY
Content-Type: multipart/form-data

请求正文​

字段类型必填说明
filefile是要分析的音频文件。

示例​

curl -X POST "https://api.genderrecognition.com/v1/audio/quality/api" \
-H "apiKey: YOUR_API_KEY" \
-F "file=@audio.wav"

响应​

{
"success": true,
"quality_score": 82,
"quality_label": "good",
"mos_score": 4.1,
"speech": {
"detected": true,
"duration": 4.8,
"ratio": 92
},
"noise": {
"level": "low",
"snr_db": 28.4
},
"volume": {
"too_quiet": false,
"rms_dbfs": -18.2
},
"clipping": {
"detected": false,
"ratio": 0
},
"audio": {
"duration": 5.2,
"sample_rate": 44100,
"channels": 1
},
"remainingRequests": 119
}

响应字段​

字段类型说明
successboolean音频处理成功时为 true。
quality_scoreinteger or null总体感知质量,范围为 0–100。未检测到语音时为 null。
quality_labelstring or nullquality_score 对应的质量等级:excellent、good、fair 或 poor。当 quality_score 为 null 时此值也为 null。
mos_scorenumber or null预测的原始平均意见分数,范围为 1–5(越高越好)。未检测到语音时为 null。
speech.detectedboolean文件中是否检测到语音。
speech.durationnumber被判定为语音的音频时长,单位为秒。
speech.ratiointeger文件中语音所占的百分比,范围为 0–100。
noise.levelstring or null背景噪声等级:low、medium 或 high。无法测量时为 null(例如没有非语音音频可供比较)。
noise.snr_dbnumber or null语音相对于背景噪声的信噪比,单位为 dB。在 noise.level 无法测量的相同情况下为 null。
volume.too_quietboolean音频过于安静,无法可靠使用时为 true。
volume.rms_dbfsnumber响度,单位为 dBFS。0 是最大音量,数值越负表示声音越小。
clipping.detectedboolean检测到削波/失真时为 true。
clipping.ratiointeger受削波影响的音频采样所占百分比,范围为 0–100。
audio.durationnumber上传文件的时长,单位为秒。
audio.sample_ratenumber上传音频的采样率,单位为 Hz。
audio.channelsnumber音频声道数(1 = 单声道,2 = 立体声)。
remainingRequestsinteger本次成功请求扣除后剩余的 API 请求数。

quality_score、speech.ratio 和 clipping.ratio 均为整数百分比。noise.snr_db 和 volume.rms_dbfs 的单位是分贝,并非百分比。

错误情况​

缺少 API 密钥​

{ "error": "API key is required" }

API 密钥无效​

{ "error": "Invalid API key" }

缺少文件​

{ "error": "No file uploaded" }

文件过大​

{ "error": "File too large" }

意外的表单字段​

如果上传文件使用的字段名不是 file(例如额外发送了一个文件),则会返回此错误:

{ "error": "Unexpected field" }

文件类型无效​

{
"error": "Invalid file type. Allowed formats: WAV, MP3, FLAC, MP4, OGG, AIFF"
}

音频转换失败​

如果文件扩展名受支持,但 ffmpeg 无法转换文件(例如文件损坏或无法读取),则会返回此错误:

{ "error": "Audio conversion failed: <details>" }

配额已用尽​

{
"error": {
"code": "RATE_LIMIT_EXCEEDED",
"message": "You have exceeded your free tier limit of 50 requests.",
"detailedMessage": "Insufficient remaining requests. Required: 1, Available: 0",
"details": {
"limit": 50,
"used": 50,
"reset_date": "2026-07-01T00:00:00.000Z"
},
"suggested_action": "Please upgrade to a premium plan to continue using the API."
}
}

处理失败​

{ "error": "Audio quality assessment failed" }

说明​

  • 后端接受 WAV、MP3、M4A、FLAC、OGG、WebM、AAC、Opus、WMA、AMR、3GP、AIFF、AU 等多种常见格式,并会在分析前将非 WAV 文件转换为 WAV。
  • 音频上传使用 file 字段,文件大小不得超过 10 MB。
  • 只有质量评估成功完成后才会扣除一次请求。