我想最佳化 Angular 應用程式的文字轉語音部分以節省 AI Token。當使用者第一次點擊音訊按鈕時,應用程式會將請求連同表單資料(場景、情緒、語音與語速)發送至 Firebase AI Logic。在後續的點擊中,即使表單資料未發生任何變更,系統仍會發送新的請求。這些請求不僅浪費資源,還會產生不必要的雲端成本。
當表單資料自上次語音產生以來未曾變更時,輸出的語音是完全相同的,因此我們可以安全地略過呼叫 AI。相反地,我們可以透過重複使用 HTML <audio> 元件中已快取的音訊來進行最佳化,讓使用者能夠即時重播與聆聽。
接下來,我將展示如何快取最後一次的音訊資料,並檢查何時該發送 AI 請求以及何時略過。
export interface GeneratedAudioRecord {
url: string;
prompt: string;
voice: string;
}
我們定義了一個介面,將提示詞、語音與音訊 URL 一起快取為不可分割的原子記錄。
在 TextToSpeechViewService 中,我們宣告一個型別為 GeneratedAudioRecord 的新 signal,並將 audioUrl 轉換為 computed signal:
readonly #activeAudio = signal<GeneratedAudioRecord | undefined>(undefined);
audioUrl = computed(() => this.#activeAudio()?.url);
activeAudio = this.#activeAudio.asReadonly();
我們定義了 setGeneratedAudioRecord 來撤銷先前的 URL 以防止記憶體流失,並快取新的有效提示詞、語音與 URL:
private setGeneratedAudioRecord(blob: Blob, prompt: string, voice: string) {
revokeBlobURL(this.#activeAudio()?.url);
const url = URL.createObjectURL(blob);
this.#activeAudio.set({ url, prompt, voice });
}
我們也定義了 clearAudio(),以便在服務銷毀時清理 Blob URL:
readonly #destroyRef$ = inject(DestroyRef);
constructor() {
this.#destroyRef$.onDestroy(() => this.clearAudio());
}
private clearAudio() {
revokeBlobURL(this.#activeAudio()?.url);
this.#activeAudio.set(undefined);
}
先前,generateSpeech 總是在發送請求前或發生錯誤時清除音訊。在新的版本中,它將快取的責任委派給 handleSync 與 handleStream。如果發送到 Firebase AI Logic 的新請求失敗,先前的音訊將被保留,以便使用者可以繼續重播最後一次有效的音訊產生。
修改前:
async generateSpeech(mode: GenerateSpeechMode, promptArgs: FactConfig) {
if (!promptArgs.fact || this.#loadingMode() !== 'idle') {
return;
}
revokeBlobURL(this.#audioUrl());
this.#audioUrl.set(undefined);
try {
this.#loadingMode.set(mode);
/* ... switch statement ... */
} catch (e) {
revokeBlobURL(this.#audioUrl());
this.#audioUrl.set(undefined);
/* ... throw error ... */
} finally {
this.#loadingMode.set('idle');
}
}
修改後:
async generateSpeech(mode: GenerateSpeechMode, promptArgs: FactConfig) {
if (!promptArgs.fact || this.#loadingMode() !== 'idle') {
return;
}
try {
this.#loadingMode.set(mode);
/* ... switch statement ... */
} catch (e) {
/* ... throw error ... */
} finally {
this.#loadingMode.set('idle');
}
}
當 Firebase AI Logic 成功返回合成音訊時,handleSync 會快取提示詞、語音與 URL:
修改前:
private async handleSync(promptArgs: FactConfig) {
try {
const speechService = await this.#asyncSpeechService();
const blob = await speechService.synthesize({ text: promptArgs.prompt, voice: promptArgs.voice });
this.setAudioUrl(blob)
} catch (e) {
this.handlePlaybackError(e);
throw e;
}
}
修改後:
private async handleSync(promptArgs: FactConfig) {
try {
const speechService = await this.#asyncSpeechService();
const blob = await speechService.synthesize({ text: promptArgs.prompt, voice: promptArgs.voice });
this.setGeneratedAudioRecord(blob, promptArgs.prompt, promptArgs.voice);
} catch (e) {
this.handlePlaybackError(e);
throw e;
}
}
setGeneratedAudioRecord 取代了 setAudioUrl 來更新 #activeAudio。
快取行為在 stream 與 web_audio_api 之間有所不同:
stream 模式下,一旦所有資料區塊都被記錄下來,提示詞、語音與 URL 就會被快取。web_audio_api 模式下,語音是即時串流('play and forget'),因此任何先前的靜態音訊都會透過 this.clearAudio() 清除。修改前:
private async handleStream({ prompt, voice, shouldWait = false }: FactConfig) {
try {
/* ...streaming logic... */
if (shouldWait && !abortController.signal.aborted && rawChunks.length > 0) {
this.setAudioUrl(toWavBlob(rawChunks, DEFAULT_AUDIO_TYPE));
}
} catch (e) {
if (!abortController.signal.aborted) {
this.handlePlaybackError(e);
throw e;
}
}
}
修改後:
private async handleStream({ prompt, voice, shouldWait = false }: FactConfig) {
try {
/* ...streaming logic... */
if (abortController.signal.aborted) {
return;
}
if (shouldWait && rawChunks.length > 0) {
this.setGeneratedAudioRecord(toWavBlob(rawChunks, DEFAULT_AUDIO_TYPE), prompt, voice);
} else {
this.clearAudio();
}
} catch (e) {
if (!abortController.signal.aborted) {
this.handlePlaybackError(e);
throw e;
}
}
}
handlePlaybackError 會記錄錯誤並停止當前播放。然而,現有的音訊 URL 不再被撤銷,讓使用者可以重播並聆聽先前有效的音訊:
修改前:
private async handlePlaybackError(e: unknown) {
console.error('Streaming playback failed:', e);
const audioPlayerService = await this.#asyncAudioPlayerService();
audioPlayerService.stopAll();
revokeBlobURL(this.#audioUrl());
}
修改後:
private async handlePlaybackError(e: unknown) {
console.error('Streaming playback failed:', e);
const audioPlayerService = await this.#asyncAudioPlayerService();
audioPlayerService.stopAll();
}
至此,setAudioUrl 已不再使用且可以刪除。TextToSpeechViewService 將唯讀的 activeAudio signal 公開給 TextToSpeechComponent,以便與當前的表單資料進行比對。當表單資料自上次產生後未曾變更時,系統將略過請求以節省 AI Token 並消除延遲。
我們建構了一個 computed signal 來檢查目前的表單輸入是否與快取的音訊記錄相符:
isGeneratedForCurrentInput = computed(() => {
const activeAudio = this.speechService.activeAudio();
if (!activeAudio) {
return false;
}
const trimmedVoice = activeAudio.voice.trim().toLowerCase();
const trimmedPrompt = activeAudio.prompt.trim().toLowerCase();
return (
trimmedPrompt === this.audioPrompt().trim().toLowerCase() && trimmedVoice === this.voice().trim().toLowerCase()
);
});
當使用者點擊按鈕以合成語音時,系統會評估 isGeneratedForCurrentInput。如果其值為 true(且模式不是 web_audio_api),我們將略過對 Firebase AI Logic 的請求。否則,將發送新的合成請求:
async generateSpeech(mode: GenerateSpeechMode) {
try {
const fact = this.interestingFact();
if (!fact || (mode !== 'web_audio_api' && this.isGeneratedForCurrentInput())) {
return;
}
this.ttsError.set('');
await this.speechService.generateSpeech(mode, {
prompt: this.audioPrompt(),
voice: this.voice(),
fact,
});
}
}
當第一次點擊按鈕時,請求會發送至 Firebase AI Logic 並建立音訊 URL。使用者會看到 HTML <audio> 元素並可重複播放音訊。當使用者在表單資料未變更的情況下再次點擊按鈕時,不會發送任何請求,使用者繼續使用現有的音訊。當使用者輸入新的表單資料時,Firebase AI Logic 則會合成新的語音。
由於在新的合成嘗試期間會保留先前的 Blob URL,我們將範本條件更新為 @if (!isLoading() && audioUrl(); as url)。這可以防止舊的 <audio> 播放器在新的請求正在載入時顯示或與即時音訊發生衝突。
第 31 天的內容就到此結束!明天,我將在發送行內 Base64 資料至 Firebase AI Logic 進行影像分析之前,先縮小所選圖片的尺寸。使用縮小尺寸的圖片來產生替代文字、推薦內容與標籤,將能大幅節省輸入 Token。