跳到主要内容

说话人验证 API

比较参考音频和候选音频,获取同一说话人判定、相似度分数以及两个文件的音频元数据。

基础 URL​

https://api.genderrecognition.com

端点​

POST /v1/speaker-verification/api

请求标头​

apiKey: YOUR_API_KEY
Content-Type: multipart/form-data

请求正文​

字段类型必填说明
referencefile是用作比对基准的已知说话人语音。
candidatefile是要与参考语音进行比对的音频。

示例​

curl -X POST "https://api.genderrecognition.com/v1/speaker-verification/api" \
-H "apiKey: YOUR_API_KEY" \
-F "file=@audio.wav"

响应​

{
"success": true,
"same_speaker": true,
"similarity": 87,
"threshold": 70,
"reference_audio": {
"duration": 4.8,
"sample_rate": 16000
},
"candidate_audio": {
"duration": 5.1,
"sample_rate": 16000
},
"remainingRequests": 119
}

响应字段​

字段类型说明
successboolean两个片段均成功处理时为 true。
same_speakerboolean当 similarity 大于或等于 threshold 时为 true,表示判定两段音频来自同一说话人。
similarityinteger两段语音的匹配程度,范围为 0–100。数值越高,相似度越高。
thresholdinteger用于判定 same_speaker 的相似度阈值,范围为 0–100。similarity 达到或超过此值时视为匹配。
reference_audio.durationnumber解码后用于比对的参考音频时长,单位为秒。
reference_audio.sample_ratenumber比对前参考音频解码后的采样率,单位为 Hz。
candidate_audio.durationnumber解码后用于比对的候选音频时长,单位为秒。
candidate_audio.sample_ratenumber比对前候选音频解码后的采样率,单位为 Hz。
remainingRequestsinteger本次成功请求扣除后剩余的 API 请求数。

similarity 和 threshold 均为整数百分比,并非原始分数。

错误情况​

缺少 API 密钥​

{ "error": "API key is required" }

API 密钥无效​

{ "error": "Invalid API key" }

缺少文件​

{ "error": "Both reference and candidate files are required" }

文件过大​

{ "error": "File too large" }

意外的表单字段​

如果文件上传时使用的字段名不是 reference 或 candidate(例如使用 file,或额外发送文件),则会返回此错误:

{ "error": "Unexpected field" }

文件类型无效​

{
"error": "Invalid file type. Allowed formats: WAV, MP3, FLAC, MP4, OGG, AIFF"
}

音频转换失败​

如果文件扩展名受支持,但 ffmpeg 无法转换文件(例如文件损坏或无法读取),则会返回此错误:

{ "error": "Audio conversion failed: <details>" }

配额已用尽​

{
"error": {
"code": "RATE_LIMIT_EXCEEDED",
"message": "You have exceeded your free tier limit of 50 requests.",
"detailedMessage": "Insufficient remaining requests. Required: 1, Available: 0",
"details": {
"limit": 50,
"used": 50,
"reset_date": "2026-07-01T00:00:00.000Z"
},
"suggested_action": "Please upgrade to a premium plan to continue using the API."
}
}

处理失败​

{ "error": "Speaker verification failed" }

说明​

  • 后端接受 WAV、MP3、M4A、FLAC、OGG、WebM、AAC、Opus、WMA、AMR、3GP、AIFF、AU 等多种常见格式,并会在分析前将非 WAV 文件转换为 WAV。
  • 两个上传文件分别使用 reference 和 candidate 字段,且大小都不得超过 10 MB。
  • 参考音频和候选音频会分别处理,因此时长、采样率和格式可以不同。
  • 只有验证成功完成后才会扣除一次请求。