Fields
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
Duration | *float64 | :heavy_minus_sign: | Duration of the input audio in seconds, present when response_format is verbose_json | 9.2 |
Language | *string | :heavy_minus_sign: | Detected or forced language, present when response_format is verbose_json | english |
Segments | []components.STTSegment | :heavy_minus_sign: | Timestamped transcript segments, present when response_format is verbose_json | |
Task | *string | :heavy_minus_sign: | The task performed, present when response_format is verbose_json | transcribe |
Text | string | :heavy_check_mark: | The transcribed text | Hello, this is a test of OpenAI speech-to-text transcription. The weather is sunny today and the temperature is around 72 degrees. |
Usage | *components.STTUsage | :heavy_minus_sign: | Aggregated usage statistics for the request | { “cost”: 0.000508, “input_tokens”: 83, “output_tokens”: 30, “seconds”: 9.2, “total_tokens”: 113 } |
Words | []components.STTWord | :heavy_minus_sign: | Timestamped words, present when the provider returns word-level timestamps |