在昨天的 [Day 07] 中,我們成功在 Next.js 後端架設了第一支調用 Gemini 1.5 Flash 的 AI API Route。然而在真實的 SaaS 商業場景中,知識工作者很少會花時間把 50 頁的財報或技術白皮書手動複製、貼成純文字餵給 AI。
用戶最直覺的操作是:「把整份 PDF 檔案拖進來,30 秒內給我重點摘要與社群簡報!」
在過去,要實現這個需求是一場工程噩夢。今天我們就來聊聊為什麼 Gemini 的 百萬級 Token 上下文視窗(Long Context Window) 能對傳統文件檢索架構進行「降維打擊」,並用代碼實作 OmniVibe AI 的長篇 PDF 解析引擎!
在 Gemini 1.5 問世之前,業界處理長文件的標準架構是 RAG(檢索增強生成,Retrieval-Augmented Generation):
graph LR
subgraph 傳統 RAG 繁複管線
A[PDF 原始文件] --> B[文字萃取 PyPDF/OCR]
B --> C[切塊 Chunking 500字]
C --> D[計算 Embedding]
D --> E[存入向量資料庫]
E --> F[向量相似度檢索 Top-K]
F --> G[拼裝碎片上下文]
G --> H[LLM 產出回覆]
end
Gemini 1.5 原生具備 100 萬至 200 萬 Token 的巨量上下文視窗(相當於約 1 小時影片、11 小時音訊或 70 萬字小說)。
我們不再需要切塊,更不需要向量資料庫! 整份 PDF 檔案以原生格式直通模型神經網絡,Gemini 能夠在毫秒級內全局縱覽全文,精確關聯首尾邏輯,甚至連 PDF 內的排版佈局與圖表結構都能直接理解。
在 Google Gen AI SDK 中,餵入 PDF 檔案主要有兩種方式:
| 方案 | 適用場景 | 限制 | 實作方式 |
|---|---|---|---|
| Inline Data (Base64) | 適用於小於 20MB 的小型輕量檔案 | 請求 Payload 受限,過大會導致連線逾時 | 轉成 Base64 字串後放入 inlineData |
| Google AI File API | 生產級 SaaS 首選! 支援高達 2GB 的大型檔案 | 檔案暫存於 Google 伺服器 48 小時,需做生命週期管理 | 使用 GoogleAIFileManager 上傳取得 URI |
為了打造抗壓、穩健的 OmniVibe AI,我們採用 Google AI File API 來處理使用者的 PDF 上傳請求!
我們將在 Next.js 中新增一支專門處理多檔案表單(multipart/form-data)的 API 端點:src/app/api/ai/transform-pdf/route.ts。
除了 @google/generative-ai 之外,我們需要使用官方提供的伺服器端檔案管理工具:
npm install @google/generative-ai
(註:GoogleAIFileManager 已內建在 @google/generative-ai/server 子模組中)
src/app/api/ai/transform-pdf/route.ts)以下代碼展示了完整的工業級流程:
/tmp 目錄。GoogleAIFileManager 推送至 Google File API。gemini-1.5-flash 進行深度全文多模態解析。// src/app/api/ai/transform-pdf/route.ts
import { NextRequest, NextResponse } from 'next/server';
import { GoogleGenerativeAI } from '@google/generative-ai';
import { GoogleAIFileManager } from '@google/generative-ai/server';
import { OMNIVIBE_SYSTEM_INSTRUCTION } from '@/lib/gemini/prompts';
import { writeFile, unlink } from 'fs/promises';
import path from 'path';
import os from 'os';
const apiKey = process.env.GEMINI_API_KEY || '';
const genAI = new GoogleGenerativeAI(apiKey);
const fileManager = new GoogleAIFileManager(apiKey);
export async function POST(req: NextRequest) {
let tempFilePath: string | null = null;
let uploadedFileResource: any = null;
try {
// 1. 解析 multipart/form-data
const formData = await req.formData();
const file = formData.get('file') as File | null;
if (!file) {
return NextResponse.json(
{ error: 'Bad Request', message: '請上傳 PDF 檔案' },
{ status: 400 }
);
}
if (file.type !== 'application/pdf') {
return NextResponse.json(
{ error: 'Bad Request', message: '僅支援 PDF 格式文件' },
{ status: 400 }
);
}
// 2. 將上傳檔案寫入伺服器暫存目錄 (/tmp)
const arrayBuffer = await file.arrayBuffer();
const buffer = Buffer.from(arrayBuffer);
const fileName = `upload_${Date.now()}_${file.name}`;
tempFilePath = path.join(os.tmpdir(), fileName);
await writeFile(tempFilePath, buffer);
// 3. 上傳檔案至 Google AI File API
uploadedFileResource = await fileManager.uploadFile(tempFilePath, {
mimeType: 'application/pdf',
displayName: file.name,
});
console.log(`[Google File API] 上傳成功: ${uploadedFileResource.file.uri}`);
// 4. 初始化 Gemini 1.5 Flash 模型
const model = genAI.getGenerativeModel({
model: 'gemini-1.5-flash',
systemInstruction: OMNIVIBE_SYSTEM_INSTRUCTION,
generationConfig: {
temperature: 0.3, // 閱讀長篇正式報告時,略微調低以增強事實精準度
maxOutputTokens: 4096,
},
});
// 5. 構建多模態 Prompt 請求
const prompt = `
請通讀這份完整的 PDF 報告,並依據你的系統人設,產出高價值的提煉成果:
1. 【核心摘要】:以金字塔結構提煉出 3~5 點全局洞見,務必保留文件中的關鍵數據與重要結論。
2. 【社群爆款矩陣】:產出一篇適合發在 LinkedIn 或 Threads 上的乾貨長文。
3. 【關鍵數據表】:將文中散落的核心指標整理為 Markdown 表格。
`;
const result = await model.generateContent([
{
fileData: {
mimeType: uploadedFileResource.file.mimeType,
fileUri: uploadedFileResource.file.uri,
},
},
{ text: prompt },
]);
const response = await result.response;
const outputText = response.text();
// 6. 回傳成果與用量監控
return NextResponse.json({
success: true,
data: {
rawOutput: outputText,
usageMetadata: response.usageMetadata,
},
fileInfo: {
name: file.name,
sizeBytes: file.size,
},
});
} catch (error: any) {
console.error('[PDF Transform Error]:', error);
return NextResponse.json(
{ error: 'Internal Server Error', message: error.message || 'PDF 解析失敗' },
{ status: 500 }
);
} finally {
// 7. 清理機制:無論成功失敗,均清理伺服器暫存與 Google 雲端臨時檔
if (tempFilePath) {
await unlink(tempFilePath).catch((err) =>
console.error('刪除本機暫存失敗:', err)
);
}
if (uploadedFileResource?.file?.name) {
await fileManager
.deleteFile(uploadedFileResource.file.name)
.catch((err) => console.error('清理 Google File API 暫存失敗:', err));
}
}
}
我們拿一份近 30 頁、包含大量圖表與排版的「2026 AI 產業趨勢白皮書」PDF 進行測試:
curl -X POST http://localhost:3000/api/ai/transform-pdf \
-F "file=@/path/to/AI_Trends_2026_Report.pdf"
{
"success": true,
"data": {
"rawOutput": "### 💡 核心洞見 (Core Takeaways)\n1. **長上下文普及引發架構顛覆**:報告第 12 頁數據指出,超過 64% 的企業開始簡化向量檢索管線...\n\n### 📊 關鍵數據對照表\n| 技術指標 | 2025 年水準 | 2026 年預測 | 成長幅度 |\n| :--- | :--- | :--- | :--- |\n| Context Window | 128k Tokens | 2M+ Tokens | 15.6x |\n| 推論成本/百萬Token | $1.50 | $0.15 | -90% |",
"usageMetadata": {
"promptTokenCount": 42810,
"candidatesTokenCount": 850,
"totalTokenCount": 43660
}
},
"fileInfo": {
"name": "AI_Trends_2026_Report.pdf",
"sizeBytes": 8421050
}
}
可以看到:
今天我們見證了大模型時代的長文本革命:
既然 PDF 都能秒級消化,那麼難度更高的「長影片」與「Podcast 音訊」呢?
👉 明天(Day 09),我們將進入【多模態震撼篇】:影音不求人!我們不再需要任何額外的語音轉文字(STT)工具,看我們如何直接將 MP4 影片與 MP3 音訊餵入 Gemini,秒產具備時間軸標記的分段精華!
精彩內容不容錯過,我們明天見!🔥