在昨天的 [Day 08] 中,我們見證了 Gemini 原生長文本上下文對傳統 RAG 架構的震撼打擊。今天,我們要來攻克內容創作者與知識工作者日常最耗時、也最棘手的痛點 —— 影音內容的精煉與再創作。
如果你曾經嘗試打造過 Podcast 總結工具或 YouTube 影片精華助手,你的後端架構大概率長這樣:
graph LR
subgraph 傳統影音處理拼裝車
A[MP4/MP3 檔案] --> B[FFmpeg 抽取音軌]
B --> C[OpenAI Whisper / STT 轉錄逐字稿]
C --> D[文字斷詞與清洗]
D --> E[文字丟給 LLM 進行摘要]
end
這套「拼裝車」架構在過去兩年是主流,但它有三個致命傷:
今天,Google Gemini 1.5 將徹底顛覆這個遊戲規則。
Gemini 不是「先在內部轉逐字稿再看文字」,它的神經網絡具備真正的原生多模態感官(Native Multimodality):
這意味著:Gemini 是一邊「看著螢幕畫面」,一邊「聽著說話內容」在進行邏輯推理。
在昨天處理 PDF 時,檔案上傳後幾乎瞬間就能調用;但大型影音檔案不同。
當你將幾百 MB 的 MP4 影片上傳到 Google AI File API 時,Google 雲端需要數秒到十數秒的時間對影格進行切片與預處理。此時檔案狀態會是 PROCESSING,若直接呼叫模型會拋出例外。
因此,生產級的影音處理後端必須實作 「非同步狀態輪詢(Active State Polling)」:
sequenceDiagram
participant Client as 前端客戶端
participant NextServer as Next.js API Route
participant GoogleAPI as Google AI File API
participant Gemini as Gemini 1.5 Flash
Client->>NextServer: POST 上傳影片 (multipart/form-data)
NextServer->>GoogleAPI: fileManager.uploadFile()
GoogleAPI-->>NextServer: 回傳 File Resource (State: PROCESSING)
loop 每隔 3 秒檢查一次狀態 (最長等待 60 秒)
NextServer->>GoogleAPI: fileManager.getFile(fileName)
GoogleAPI-->>NextServer: 回傳當前 State
end
Note over NextServer,GoogleAPI: 檔案狀態變為 ACTIVE
NextServer->>Gemini: model.generateContent([fileData, prompt])
Gemini-->>NextServer: 產出帶有時間戳記的結構化精華
NextServer->>GoogleAPI: fileManager.deleteFile(fileName) [清理臨時檔]
NextServer-->>Client: 回傳成功 JSON
我們在 src/app/api/ai/transform-media/route.ts 建立專門處理影音上傳的端點。
在 src/lib/gemini/file-helper.ts 建立輪詢檢查邏輯:
// src/lib/gemini/file-helper.ts
import { GoogleAIFileManager, FileState } from '@google/generative-ai/server';
/**
* 輪詢等待 Google AI File API 處理完成 (狀態變為 ACTIVE)
*/
export async function waitForFileActive(
fileManager: GoogleAIFileManager,
fileName: string,
maxWaitMs: number = 60000,
pollIntervalMs: number = 3000
) {
const startTime = Date.now();
while (Date.now() - startTime < maxWaitMs) {
const file = await fileManager.getFile(fileName);
if (file.state === FileState.ACTIVE) {
console.log(`[Google File API] 檔案 ${fileName} 已就緒 (ACTIVE)`);
return file;
}
if (file.state === FileState.FAILED) {
throw new Error(`[Google File API] 影音檔案處理失敗: ${file.error?.message || '未知錯誤'}`);
}
console.log(`[Google File API] 檔案處理中 (${file.state})... 等待 ${pollIntervalMs / 1000} 秒`);
await new Promise((resolve) => setTimeout(resolve, pollIntervalMs));
}
throw new Error(`[Google File API] 等待檔案處理逾時 (超過 ${maxWaitMs / 1000} 秒)`);
}
src/app/api/ai/transform-media/route.ts)這支 API 支援音訊(audio/*)與影片(video/*),並指示 Gemini 透過時間軸(Timestamps)精確捕捉關鍵畫面與言論:
// src/app/api/ai/transform-media/route.ts
import { NextRequest, NextResponse } from 'next/server';
import { GoogleGenerativeAI } from '@google/generative-ai';
import { GoogleAIFileManager } from '@google/generative-ai/server';
import { waitForFileActive } from '@/lib/gemini/file-helper';
import { OMNIVIBE_SYSTEM_INSTRUCTION } from '@/lib/gemini/prompts';
import { writeFile, unlink } from 'fs/promises';
import path from 'path';
import os from 'os';
const apiKey = process.env.GEMINI_API_KEY || '';
const genAI = new GoogleGenerativeAI(apiKey);
const fileManager = new GoogleAIFileManager(apiKey);
export async function POST(req: NextRequest) {
let tempFilePath: string | null = null;
let uploadedFileResource: any = null;
try {
const formData = await req.formData();
const file = formData.get('file') as File | null;
if (!file) {
return NextResponse.json({ error: '請提供多媒體檔案' }, { status: 400 });
}
// 支援 MP4, MOV, MP3, WAV, M4A
const isVideo = file.type.startsWith('video/');
const isAudio = file.type.startsWith('audio/');
if (!isVideo && !isAudio) {
return NextResponse.json(
{ error: '不支援的檔案格式,請上傳影片或音訊檔案' },
{ status: 400 }
);
}
// 1. 寫入本地臨時磁碟
const arrayBuffer = await file.arrayBuffer();
const buffer = Buffer.from(arrayBuffer);
const fileName = `media_${Date.now()}_${file.name}`;
tempFilePath = path.join(os.tmpdir(), fileName);
await writeFile(tempFilePath, buffer);
// 2. 上傳至 Google AI File API
uploadedFileResource = await fileManager.uploadFile(tempFilePath, {
mimeType: file.type,
displayName: file.name,
});
console.log(`[Google File API] 上傳完成,File Name: ${uploadedFileResource.file.name}`);
// 3. 如果是影片,需要輪詢等待 Google 完成切片處理
if (isVideo) {
await waitForFileActive(fileManager, uploadedFileResource.file.name);
}
// 4. 呼叫 Gemini 1.5 Flash
const model = genAI.getGenerativeModel({
model: 'gemini-1.5-flash',
systemInstruction: OMNIVIBE_SYSTEM_INSTRUCTION,
generationConfig: {
temperature: 0.4,
maxOutputTokens: 4096,
},
});
// 5. 針對影音特化的 Prompt:要求包含精確的時間軸與視覺畫面描述
const mediaPrompt = `
這是一份使用者上傳的${isVideo ? '影音錄像' : '音訊錄音'}。請同時結合聽覺對話與視覺畫面(若為影片),為內容創作者產出以下高價值資產:
1. 【分段時間戳記 (Key Chapters)】:
- 格式:[MM:SS] 章節標題 - 核心內容簡述(若有投影片或畫面特徵,請一併指出)。
2. 【爆款短影音剪輯建議 (Short-form Clips)】:
- 挑出 2~3 個最適合剪成 60 秒短影音(Reels / TikTok)的高光片段。
- 標明開始與結束時間戳記(例:[02:15 - 03:10])。
- 提供「開頭黃金鉤子 (Hook)」與畫面剪輯建議。
3. 【金句摘錄 (Punchlines)】:精確記錄講者說過的 3 句最具傳播力的原話與出現時間點。
`;
const result = await model.generateContent([
{
fileData: {
mimeType: uploadedFileResource.file.mimeType,
fileUri: uploadedFileResource.file.uri,
},
},
{ text: mediaPrompt },
]);
const response = await result.response;
return NextResponse.json({
success: true,
data: {
rawOutput: response.text(),
usageMetadata: response.usageMetadata,
},
mediaType: isVideo ? 'video' : 'audio',
});
} catch (error: any) {
console.error('[Media Transform Error]:', error);
return NextResponse.json(
{ error: '伺服器處理影音失敗', message: error.message },
{ status: 500 }
);
} finally {
// 6. 安全清理機制
if (tempFilePath) {
await unlink(tempFilePath).catch(() => {});
}
if (uploadedFileResource?.file?.name) {
await fileManager.deleteFile(uploadedFileResource.file.name).catch(() => {});
}
}
}
我們將一段實錄的技術發表會 MP4 檔案(包含講者投影片展示與現場操作 Demo)發送至端點:
curl -X POST http://localhost:3000/api/ai/transform-media \
-F "file=@/path/to/product_demo_8min.mp4"
### ⏱️ 分段時間戳記 (Key Chapters)
* ** 破題開場**:講者展示當前 AI SaaS 面臨的架構碎片化痛點,螢幕投影出傳統架構圖。
* ** 核心架構發布**:畫面切換至 OmniVibe 全棧藍圖,講者特別指著右方 Google AI 生態圈模組進行說明。
* ** 實機 Live Demo 翻車與應變**:現場上傳 100MB 影片實測,講者幽默化解等待時間,隨後 5 秒內展示出生成結果。
* ** 總結與 Q&A**:公布開源倉庫連結與未來功能 Roadmap。
---
### 🎬 爆款短影音剪輯建議 (Viral Clips)
#### 推薦片段 1:【架構批判】
* **時間軸**:[00:45 - 01:40] (長度 55 秒)
* **前 3 秒黃金鉤子 (Hook)**:「還在用 5 種不同的第三方服務拼裝你的 AI 後端?你正在慢性自殺!」
* **視覺與剪輯提示**:
* 在 處將鏡頭 Zoom in 講者激動的手勢。
* 背景音效在講出「帳單爆表」時下重音。
---
### 💬 現場金句原話 (Punchlines)
1. 「寫代碼不是你的資產,解決問題的商業閉環才是。」 ——
2. 「不要在沒人買你的產品之前,寫超過一行多餘的邏輯。」 ——
更不可思議的是:整個過程完全沒有用到任何外部語音辨識模型,Gemini 連影片畫面上投影片的小字與講者的手勢動作都精確記錄了下來!
今天我們解鎖了全棧 AI 開發的最高殿堂:
然而,當我們有了這麼強大的輸出內容,前端要怎麼呈現?
如果 AI 有時候回傳 Markdown、有時候漏了冒號、有時候格式微調,前端畫面就會崩潰給你看。我們需要讓 AI 回傳「100% 絕對穩定的純 JSON 格式」!
👉 明天(Day 10),我們將進入【結構化資料篇】:實戰 Structured Outputs(JSON Schema)!看我們如何用代碼給 Gemini 套上緊箍咒,保證回傳格式分毫不差,前端元件渲染永不白屏!
我們明天見!🔥