iT邦幫忙

2026 iThome 鐵人賽

DAY 10
0
AI Engineering

從行程到應援:30 天做出追星實用工具箱系列 第 10 篇

Day 10:OpenAI Whisper API 串接與 HTTP/WebSocket 傳輸機制

  • 分享至 

  • xImage
  •  

一、今日目標

  1. 理解 WebSocket 串流機制:比較 HTTP 與 WebSocket 的差異與應用場景。
  2. HTTP 傳輸實作:將錄好的音訊檔發送到 OpenAI Whisper轉換成文字。
  3. 整合 UI:在畫面上顯示語音轉文字的結果。

二、今日任務

HTTP vs. WebSocket 傳輸機制的比較

1. HTTP (REST API) =「寄信 / 叫外送」

最傳統的網路溝通方式,遵循一問一答的規則,你沒發問,伺服器就絕對不會理你。

  • 運作模式:你發送一個請求,伺服器收到後處理,最後回傳結果,這回合就結束了。
  • 生活比喻:就像你用 Uber Eats 點餐
    1. 你點完餐送出訂單(發送 HTTP 請求)
    2. 外送員把餐點送到你家(伺服器回傳 Response)
    3. 交易完成!如果還想吃別的,你必須「重新下一次單」
  • 語音應用:你把整段錄好的 10 秒語音檔一次傳過去,OpenAI 處理完後,一次把整段文字回傳給你。

2. WebSocket (串流) =「打電話 / 講對講機」

為了「即時互動」而生的技術,它會在雙方之間建立一條持久不中斷的專用通道。

  • 運作模式:只要通道一建立,雙方隨時都可以把資料像流水一樣源源不絕地丟給對方,不需要每次都重新連線。
  • 生活比喻:就像你跟朋友打電話。
    1. 電話一接通,通道就建立了。
    2. 你可以邊想邊講,對方也可以邊聽邊回應,雙方隨時都能即時互相插話、對話。
    3. 直到有人掛斷電話為止,通道才會關閉。
  • 語音應用:你嘴巴講話的同時,手機就把一小塊一小塊的語音封包即時丟過去,OpenAI 邊聽邊把辨識出的文字「一個字一個字」吐回你的螢幕上。
特性 HTTP (REST API) WebSocket (串流)
運作方式 錄音完整結束 ➔ 存檔 ➔ 一次性上傳 ➔ 等待回應 邊錄音 ➔ 邊送 ➔ AI 邊辨識邊吐字
延遲度 高 低
適用場景 語音輸入法、語音訊息、會議記錄轉寫 即時 AI 語音通話
實作難度 簡單 難

Step 1: 安裝與設定 HTTP 套件

在 pubspec.yaml 加入 http 與 path_provider(用於存暫存檔):

dependencies:
	http: ^1.2.1
  path_provider: ^2.1.2

執行 flutter pub get

Step 2: 實作 Whisper Service

建立 lib/services/whisper_service.dart,負責將音訊檔案發送給 OpenAI:

import 'dart:io';
import 'package:http/http.dart' as http;
import 'package:flutter_dotenv/flutter_dotenv.dart';
import 'dart:convert';

class WhisperService {
  final String _baseUrl = 'https://api.openai.com/v1/audio/transcriptions';

  /// 將本地音訊檔案上傳至 OpenAI Whisper API 轉換為文字
  Future<String?> transcribeAudio(File audioFile) async {
    final apiKey = dotenv.env['OPENAI_API_KEY'];
    
    if (apiKey == null || apiKey.isEmpty) {
      print('❌ 錯誤:未設定 OPENAI_API_KEY');
      return null;
    }

    try {
      // 1. 建立請求
      var request = http.MultipartRequest('POST', Uri.parse(_baseUrl));

      // 2. 設定 Request Headers
      request.headers.addAll({
        'Authorization': 'Bearer $apiKey',
      });

      // 3. 設定表單參數
      request.fields['model'] = 'whisper-1';
      request.fields['language'] = 'zh'; // 指定中文辨識,可以提升準確度(選填)

      // 4. 附加音訊檔案
      request.files.add(
        await http.MultipartFile.fromPath(
          'file',
          audioFile.path,
        ),
      );

      print('🚀 正在發送音訊至 Whisper API...');
      var streamedResponse = await request.send();
      var response = await http.Response.fromStream(streamedResponse);

      if (response.statusCode == 200) {
        var jsonResponse = jsonDecode(utf8.decode(response.bodyBytes));
        String text = jsonResponse['text'];
        print('✅ 辨識成功: $text');
        return text;
      } else {
        print('❌ API 錯誤 [${response.statusCode}]: ${response.body}');
        return null;
      }
    } catch (e) {
      print('❌ 發送請求失敗: $e');
      return null;
    }
  }
}

Step 3: 更新 UI 與錄音儲存邏輯

結合 Day 9 的錄音和今天的 Whisper API,錄音結束後自動存檔並翻譯:

import 'dart:io';
import 'package:flutter/material.dart';
import 'package:path_provider/path_provider.dart';
import 'services/audio_recorder_service.dart';
import 'services/whisper_service.dart';

class AudioRecorderPage extends StatefulWidget {
  const AudioRecorderPage({super.key});

  @override
  State<AudioRecorderPage> createState() => _AudioRecorderPageState();
}

class _AudioRecorderPageState extends State<AudioRecorderPage> {
  final AudioRecorderService _recorderService = AudioRecorderService();
  final WhisperService _whisperService = WhisperService();
  
  bool _isRecording = false;
  bool _isLoading = false;
  String _transcribedText = '';
  String? _currentAudioPath;

  void _toggleRecording() async {
    if (_isRecording) {
      // 停止錄音
      await _recorderService.stopRecording();
      setState(() {
        _isRecording = false;
        _isLoading = true;
      });

      // 錄音結束後,將檔案送往 Whisper API
      if (_currentAudioPath != null) {
        final text = await _whisperService.transcribeAudio(File(_currentAudioPath!));
        setState(() {
          _transcribedText = text ?? '辨識失敗,請重試。';
          _isLoading = false;
        });
      }
    } else {
      // 開始錄音 (儲存為 mp3 或 m4a 檔以供 HTTP 上傳)
      final tempDir = await getTemporaryDirectory();
      _currentAudioPath = '${tempDir.path}/temp_audio.m4a';

      await _recorderService.startFileRecording(_currentAudioPath!);
      setState(() {
        _isRecording = true;
        _transcribedText = '';
      });
    }
  }

  @override
  Widget build(BuildContext context) {
    return Scaffold(
      appBar: AppBar(title: const Text('Day 10 - Whisper 語音轉文字')),
      body: Padding(
        padding: const EdgeInsets.all(20.0),
        child: Column(
          mainAxisAlignment: MainAxisAlignment.center,
          children: [
            Icon(
              _isRecording ? Icons.mic : Icons.mic_none,
              size: 80,
              color: _isRecording ? Colors.red : Colors.grey,
            ),
            const SizedBox(height: 20),
            ElevatedButton(
              onPressed: _isLoading ? null : _toggleRecording,
              style: ElevatedButton.styleFrom(
                backgroundColor: _isRecording ? Colors.red : Colors.blue,
                padding: const EdgeInsets.symmetric(horizontal: 32, vertical: 16),
              ),
              child: Text(
                _isRecording ? '停止並轉文字' : '開始錄音',
                style: const TextStyle(color: Colors.white, fontSize: 16),
              ),
            ),
            const SizedBox(height: 30),
            if (_isLoading) ...[
              const CircularProgressIndicator(),
              const SizedBox(height: 10),
              const Text('AI 正在辨識語音中...'),
            ],
            if (_transcribedText.isNotEmpty) ...[
              const Text('辨識結果:', style: TextStyle(fontWeight: FontWeight.bold, fontSize: 16)),
              const SizedBox(height: 10),
              Container(
                padding: const EdgeInsets.all(16),
                decoration: BoxDecoration(
                  color: Colors.grey[200],
                  borderRadius: BorderRadius.circular(8),
                ),
                child: Text(_transcribedText, style: const TextStyle(fontSize: 16)),
              ),
            ]
          ],
        ),
      ),
    );
  }
}

檢查清單

  • [ ] 點擊「開始錄音」,對著手機/模擬器講一句話(例如:「測試 OpenAI Whisper API 轉文字」)。
  • [ ] 點擊「停止並轉文字」,查看偵錯主控台是否印出 🚀 正在發送... 與 ✅ 辨識成功。
  • [ ] 確認 App 畫面上成功印出你剛才講的話!

上一篇
Day 9:Flutter 音訊處理實作/錄音與語音低延遲切片
下一篇
Day 11:如何將切片音訊流暢送出並取得逐字稿
系列文
從行程到應援:30 天做出追星實用工具箱 共 16 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言