3

我的 .wav 文件长度只有 4 秒。即使经过多次重试并在云上运行它,我也会不断收到以下错误

  * upload completely sent off: 12 out of 12 bytes
  < HTTP/1.1 408 Request timed out (> 14000 ms)
  < Transfer-Encoding: chunked
  < Content-Type: text/plain
  < Server: Microsoft-IIS/8.5
  < X-MSEdge-Ref: 

有人遇到过这个问题吗?这是我的要求

  `curl -v "https://speech.platform.bing.com/recognize?
  scenarios=catsearch&appid=D4D52672-91D7-4C74-8AD8-42B1D98141A5&locale=en-  
  US&device.os=wp7&version=3.0&format=json&requestid=1d4b6030-9099-12e0-91e4-
  0800200c9a67&instanceid=1d4b6030-9099-12e0-91e5-0800200c9a68" -H 
  "Authorization: Bearer $1" -H "Content-Type: audio/wav; samplerate=8000" -- 
  data-binary $2`
4

2 回答 2

2

我也遇到了一些问题让它工作。以下 BASH 脚本“bingrec.sh”可能有助于使其更清晰;输入您的 SUBSCRIPTION_KEY 并根据需要调整 SAMPLERATE 等。正如其他人指出的那样,需要将语言环境和场景设置为支持的值,并且 instance_id 和 request_id 需要采用 GUID 格式。音频文件的长度应小于 10 秒,采样率为 8000 或 16000。此外, curl "--data-binary" 参数需要在音频文件名前加上 "@"。

#!/bin/bash
#  Usage:  ./bingrec.sh  /path/to/file 
#  Send audio file $1 through Bing speech recognition API.
#
SUBSCRIPTION_KEY=<your-key-here>
LOCALE=en-US
SCENARIOS=ulm
SAMPLERATE=8000
CODEC=audio/pcm

TARGET_FILE=$1
if [ ! -f "$TARGET_FILE" ]; then
  echo Error:  file $TARGET_FILE does not exist!
  exit 1
fi

INSTANCE_ID=`uuidgen`    # random GUID for instance
REQUEST_ID=`uuidgen`     # random GUID for request
APPID=D4D52672-91D7-4C74-8AD8-42B1D98141A5   # APPID for Bing Speechrec API, don't change
DEVICE_OS=linux          # arbitraty
FORMAT=json

AUTH_TOKEN=`curl -v -X POST "https://api.cognitive.microsoft.com/sts/v1.0/issueToken" -H "Content-type: application/x-www-form-urlencoded" -H "Content-Length: 0" -H "Ocp-Apim-Subscription-Key: ${SUBSCRIPTION_KEY}"`

curl -v -X POST "https://speech.platform.bing.com/recognize?scenarios=${SCENARIOS}&appid=${APPID}&locale=${LOCALE}&device.os=${DEVICE_OS}&version=3.0&format=${FORMAT}&instanceid=${INSTANCE_ID}&requestid=${REQUEST_ID}" -H "Authorization: Bearer ${AUTH_TOKEN}" -H "Content-type: audio/wav; codec='${CODEC}'; samplerate=${SAMPLERATE}" --data-binary @${TARGET_FILE}
于 2016-12-12T23:53:10.480 回答
0

我得到了这个工作。有几个问题。一个是使用语言环境,我将其更改为 en-IN。然后场景=ulm。这似乎成功了。我能够非常清楚地检测到语音。

于 2016-07-01T09:30:44.667 回答