iT邦幫忙

2026 iThome 鐵人賽

0

我想最佳化 Angular 應用程式的文字轉語音部分以節省 AI Token。當使用者第一次點擊音訊按鈕時,應用程式會將請求連同表單資料(場景、情緒、語音與語速)發送至 Firebase AI Logic。在後續的點擊中,即使表單資料未發生任何變更,系統仍會發送新的請求。這些請求不僅浪費資源,還會產生不必要的雲端成本。

當表單資料自上次語音產生以來未曾變更時,輸出的語音是完全相同的,因此我們可以安全地略過呼叫 AI。相反地,我們可以透過重複使用 HTML <audio> 元件中已快取的音訊來進行最佳化,讓使用者能夠即時重播與聆聽。

接下來,我將展示如何快取最後一次的音訊資料,並檢查何時該發送 AI 請求以及何時略過。

快取最後的音訊提示詞、語音與音訊 URL

定義音訊記錄介面

export interface GeneratedAudioRecord {
  url: string;
  prompt: string;
  voice: string;
}

我們定義了一個介面,將提示詞、語音與音訊 URL 一起快取為不可分割的原子記錄。

在 TextToSpeechViewService 中加入快取

在 TextToSpeechViewService 中,我們宣告一個型別為 GeneratedAudioRecord 的新 signal,並將 audioUrl 轉換為 computed signal:

readonly #activeAudio = signal<GeneratedAudioRecord | undefined>(undefined);
audioUrl = computed(() => this.#activeAudio()?.url);
activeAudio = this.#activeAudio.asReadonly();

我們定義了 setGeneratedAudioRecord 來撤銷先前的 URL 以防止記憶體流失,並快取新的有效提示詞、語音與 URL:

private setGeneratedAudioRecord(blob: Blob, prompt: string, voice: string) {
    revokeBlobURL(this.#activeAudio()?.url);
    const url = URL.createObjectURL(blob);
    this.#activeAudio.set({ url, prompt, voice });
}

我們也定義了 clearAudio(),以便在服務銷毀時清理 Blob URL:

readonly #destroyRef$ = inject(DestroyRef);

constructor() {
    this.#destroyRef$.onDestroy(() => this.clearAudio());
}

private clearAudio() {
    revokeBlobURL(this.#activeAudio()?.url);
    this.#activeAudio.set(undefined);
}

修改 TextToSpeechViewService 中的 Generate Speech

先前,generateSpeech 總是在發送請求前或發生錯誤時清除音訊。在新的版本中,它將快取的責任委派給 handleSync 與 handleStream。如果發送到 Firebase AI Logic 的新請求失敗,先前的音訊將被保留,以便使用者可以繼續重播最後一次有效的音訊產生。

修改前:

async generateSpeech(mode: GenerateSpeechMode, promptArgs: FactConfig) {
    if (!promptArgs.fact || this.#loadingMode() !== 'idle') {
      return;
    }

    revokeBlobURL(this.#audioUrl());
    this.#audioUrl.set(undefined);

    try {
      this.#loadingMode.set(mode);
      /* ... switch statement ... */
    } catch (e) {
        revokeBlobURL(this.#audioUrl());
        this.#audioUrl.set(undefined);
        /* ... throw error ... */
    } finally {
      this.#loadingMode.set('idle');
    }
}

修改後:

async generateSpeech(mode: GenerateSpeechMode, promptArgs: FactConfig) {
    if (!promptArgs.fact || this.#loadingMode() !== 'idle') {
      return;
    }

    try {
      this.#loadingMode.set(mode);
      /* ... switch statement ... */
    } catch (e) {
        /* ... throw error ... */
    } finally {
      this.#loadingMode.set('idle');
    }
}

為同步模式新增快取

當 Firebase AI Logic 成功返回合成音訊時,handleSync 會快取提示詞、語音與 URL:

修改前:

private async handleSync(promptArgs: FactConfig) {
    try {
      const speechService = await this.#asyncSpeechService();
      const blob = await speechService.synthesize({ text: promptArgs.prompt, voice: promptArgs.voice });
      this.setAudioUrl(blob)
    } catch (e) {
      this.handlePlaybackError(e);
      throw e;
    }
}

修改後:

private async handleSync(promptArgs: FactConfig) {
    try {
      const speechService = await this.#asyncSpeechService();
      const blob = await speechService.synthesize({ text: promptArgs.prompt, voice: promptArgs.voice });
      this.setGeneratedAudioRecord(blob, promptArgs.prompt, promptArgs.voice);
    } catch (e) {
      this.handlePlaybackError(e);
      throw e;
    }
}

setGeneratedAudioRecord 取代了 setAudioUrl 來更新 #activeAudio。

在串流與 Web Audio 模式中加入快取

快取行為在 stream 與 web_audio_api 之間有所不同:

  • 在 stream 模式下,一旦所有資料區塊都被記錄下來,提示詞、語音與 URL 就會被快取。
  • 在 web_audio_api 模式下,語音是即時串流('play and forget'),因此任何先前的靜態音訊都會透過 this.clearAudio() 清除。

修改前:

private async handleStream({ prompt, voice, shouldWait = false }: FactConfig) {
    try {
        /* ...streaming logic... */

        if (shouldWait && !abortController.signal.aborted && rawChunks.length > 0) {
            this.setAudioUrl(toWavBlob(rawChunks, DEFAULT_AUDIO_TYPE));
        }
    } catch (e) {
      if (!abortController.signal.aborted) {
        this.handlePlaybackError(e);
        throw e;
      }
    }
}

修改後:

private async handleStream({ prompt, voice, shouldWait = false }: FactConfig) {
    try {
        /* ...streaming logic... */

        if (abortController.signal.aborted) {
            return;
        }

        if (shouldWait && rawChunks.length > 0) {
            this.setGeneratedAudioRecord(toWavBlob(rawChunks, DEFAULT_AUDIO_TYPE), prompt, voice);
        } else {
            this.clearAudio();
        }
    } catch (e) {
      if (!abortController.signal.aborted) {
        this.handlePlaybackError(e);
        throw e;
      }
    }
}

例外狀況處理

handlePlaybackError 會記錄錯誤並停止當前播放。然而,現有的音訊 URL 不再被撤銷,讓使用者可以重播並聆聽先前有效的音訊:

修改前:

private async handlePlaybackError(e: unknown) {
    console.error('Streaming playback failed:', e);
    const audioPlayerService = await this.#asyncAudioPlayerService();
    audioPlayerService.stopAll();
    revokeBlobURL(this.#audioUrl());
}

修改後:

private async handlePlaybackError(e: unknown) {
    console.error('Streaming playback failed:', e);
    const audioPlayerService = await this.#asyncAudioPlayerService();
    audioPlayerService.stopAll();
}

至此,setAudioUrl 已不再使用且可以刪除。TextToSpeechViewService 將唯讀的 activeAudio signal 公開給 TextToSpeechComponent,以便與當前的表單資料進行比對。當表單資料自上次產生後未曾變更時,系統將略過請求以節省 AI Token 並消除延遲。

在 TextToSpeechComponent 中防範多餘的 Firebase TTS 請求

我們建構了一個 computed signal 來檢查目前的表單輸入是否與快取的音訊記錄相符:

isGeneratedForCurrentInput = computed(() => {
    const activeAudio = this.speechService.activeAudio();
    if (!activeAudio) {
      return false;
    }

    const trimmedVoice = activeAudio.voice.trim().toLowerCase();
    const trimmedPrompt = activeAudio.prompt.trim().toLowerCase();
    return (
      trimmedPrompt === this.audioPrompt().trim().toLowerCase() && trimmedVoice === this.voice().trim().toLowerCase()
    );
  });

當使用者點擊按鈕以合成語音時,系統會評估 isGeneratedForCurrentInput。如果其值為 true(且模式不是 web_audio_api),我們將略過對 Firebase AI Logic 的請求。否則,將發送新的合成請求:

async generateSpeech(mode: GenerateSpeechMode) {
    try {
      const fact = this.interestingFact();
      if (!fact || (mode !== 'web_audio_api' && this.isGeneratedForCurrentInput())) {
        return;
      }

      this.ttsError.set('');
      await this.speechService.generateSpeech(mode, {
        prompt: this.audioPrompt(),
        voice: this.voice(),
        fact,
      });
    }
}

當第一次點擊按鈕時,請求會發送至 Firebase AI Logic 並建立音訊 URL。使用者會看到 HTML <audio> 元素並可重複播放音訊。當使用者在表單資料未變更的情況下再次點擊按鈕時,不會發送任何請求,使用者繼續使用現有的音訊。當使用者輸入新的表單資料時,Firebase AI Logic 則會合成新的語音。

傳輸中音訊播放器抑制

由於在新的合成嘗試期間會保留先前的 Blob URL,我們將範本條件更新為 @if (!isLoading() && audioUrl(); as url)。這可以防止舊的 <audio> 播放器在新的請求正在載入時顯示或與即時音訊發生衝突。

第 31 天的內容就到此結束!明天,我將在發送行內 Base64 資料至 Firebase AI Logic 進行影像分析之前,先縮小所選圖片的尺寸。使用縮小尺寸的圖片來產生替代文字、推薦內容與標籤,將能大幅節省輸入 Token。

相關資源


上一篇
Day 30 - 實作主畫面並測量其效能
系列文
2026年,如何利用 Antigravity CLI、Gemini、各項技能及 MCP Server 建構基於 Firebase 的 Angular 應用 共 31 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言