在昨天的 [Day 20] 中,我們成功導入了 LLM-as-a-Judge 品質評測與 OpenTelemetry 鏈路追蹤管線,讓 OmniVibe AI 具備了量化品質與檢測 AI 幻覺的能力。
然而,隨著使用者大量湧入,我們很快面臨了全棧 AI SaaS 的終極挑戰——高昂的 Token 算力成本與回應延遲(Latency)!
當使用者上傳一部 2 小時的 Podcast 影片或 300 頁的 PDF 報告時,背後高達數十萬字(100k+ Tokens)的脈絡資料。若每次對該資產進行提煉、問答或生成腳本,都要重複傳送這 100k+ Tokens 給 Gemini 1.5,不僅每次呼叫都需要等待 3 到 5 秒,更會迅速吃光使用者的免費額度與我們的 API 預算。
今天,我們將介紹 Gemini 1.5 專為長脈絡設計的殺手級功能 Context Caching(上下文快取),並搭配 Qdrant 向量資料庫實作 Semantic Cache(語意快取),雙管齊下降低 75% 以上的 API 成本,並將重複查詢拉升至毫秒級回應(sub-50ms)!
為了達到極致的效能與成本最佳化,我們設計了雙層快取防線:
flowchart TD
A[使用者發起請求/提問] --> B{L1: Qdrant 語意快取比對}
B -- 向量相似度 >= 0.92 (Hit) --> C[直接回傳快取結果 (Latency < 50ms, Cost $0)]
B -- 快取未命中 (Miss) --> D{資產 > 32k Tokens?}
D -- 是 --> E{L2: Gemini Context Cache 存在?}
E -- 存在 (Hit) --> F[使用 cachedContent 呼叫 Gemini (75% Cost Reduction)]
E -- 不存在 (Miss) --> G[建立 Google Context Cache 物件] --> F
D -- 否 --> H[標準 Gemini 1.5 Flash 呼叫]
F --> I[取得 Gemini 生成結果]
H --> I
I --> J[將新結果寫入 L1 語意快取]
J --> K[回傳給前端使用者]
src/lib/gemini/context-cache.ts)Gemini 1.5 支援將長文字、音訊、影片等 File API 上傳的資產快取於雲端端點。我們利用 @google/generative-ai/server SDK 的 GoogleAICacheManager 來維護上下文快取:
// src/lib/gemini/context-cache.ts
import { GoogleAICacheManager } from '@google/generative-ai/server';
import { GoogleGenerativeAI } from '@google/generative-ai';
const cacheManager = new GoogleAICacheManager(process.env.GEMINI_API_KEY || '');
const genAI = new GoogleGenerativeAI(process.env.GEMINI_API_KEY || '');
/**
* 建立 Gemini 長脈絡快取 (適用於 > 32,768 Tokens 的長資產)
*/
export async function createOrGetContextCache(
assetId: string,
fileUri: string,
mimeType: string,
ttlMinutes: number = 60
) {
const cacheKey = `cache-asset-${assetId}`;
try {
// 1. 嘗試查詢是否已有現存且未過期的 Cache
const existingCaches = await cacheManager.list();
const activeCache = existingCaches.cachedContents?.find(
(c) => c.displayName === cacheKey && new Date(c.expireTime) > new Date()
);
if (activeCache) {
console.log(`[Gemini Cache] 命中現存 Context Cache: ${activeCache.name}`);
return activeCache.name;
}
console.log(`[Gemini Cache] 未命中快取,建立新的長脈絡快取 (TTL: ${ttlMinutes}m)...`);
// 2. 建立全新的 Context Cache 物件
const cache = await cacheManager.create({
model: 'models/gemini-1.5-flash-002',
displayName: cacheKey,
contents: [
{
role: 'user',
parts: [{ fileData: { fileUri, mimeType } }],
},
],
ttlSeconds: ttlMinutes * 60,
});
console.log(`[Gemini Cache] 快取建立成功!Name: ${cache.name}`);
return cache.name;
} catch (error) {
console.error('[Gemini Cache Error] 建立快取失敗,降級回標準呼叫:', error);
return null;
}
}
/**
* 使用 Context Cache 進行極速推論
*/
export async function generateWithContextCache(cacheName: string, prompt: string) {
// 使用 getGenerativeModelFromCachedContent 載入快取模型
const model = genAI.getGenerativeModelFromCachedContent({
cachedContent: { name: cacheName } as any,
model: 'gemini-1.5-flash-002',
});
const result = await model.generateContent(prompt);
return {
text: result.response.text(),
usageMetadata: result.response.usageMetadata,
};
}
src/lib/cache/semantic-cache.ts)即便有了 Context Caching,如果不同使用者或同一位使用者重複詢問相似的問題(例如:「這部影片的重點是什麼?」與「請總結這部影片的精華」),呼叫大模型依然會產生費用。
我們利用 text-embedding-004 計算 Prompt 的向量,並在 Qdrant 向量資料庫中搜尋相似度:
// src/lib/cache/semantic-cache.ts
import { QdrantClient } from '@qdrant/js-client-rest';
import { GoogleGenerativeAI } from '@google/generative-ai';
const qdrant = new QdrantClient({
url: process.env.QDRANT_URL || 'http://localhost:6333',
apiKey: process.env.QDRANT_API_KEY,
});
const genAI = new GoogleGenerativeAI(process.env.GEMINI_API_KEY || '');
const COLLECTION_NAME = 'omnivibe_semantic_cache';
/**
* 取得 Prompt 的 768 維 Embedding 向量
*/
async function getEmbedding(text: string): Promise<number[]> {
const embeddingModel = genAI.getGenerativeModel({ model: 'text-embedding-004' });
const result = await embeddingModel.embedContent(text);
return result.embedding.values;
}
/**
* 查詢 L1 語意快取
*/
export async function searchSemanticCache(assetId: string, prompt: string) {
try {
const vector = await getEmbedding(prompt);
const searchResults = await qdrant.search(COLLECTION_NAME, {
vector,
filter: {
must: [{ key: 'assetId', match: { value: assetId } }],
},
limit: 1,
score_threshold: 0.92, // 向量餘弦相似度門檻高於 0.92 視為相同意圖
});
if (searchResults.length > 0) {
console.log(`[Semantic Cache] 🎯 命中語意快取!Similarity Score: ${searchResults[0].score}`);
return searchResults[0].payload?.cachedResponse as string;
}
} catch (err) {
console.warn('[Semantic Cache Warning] 快取查詢異常:', err);
}
return null;
}
/**
* 寫入 L1 語意快取
*/
export async function saveSemanticCache(assetId: string, prompt: string, responseText: string) {
try {
const vector = await getEmbedding(prompt);
await qdrant.upsert(COLLECTION_NAME, {
points: [
{
id: crypto.randomUUID(),
vector,
payload: {
assetId,
prompt,
cachedResponse: responseText,
createdAt: new Date().toISOString(),
},
},
],
});
} catch (err) {
console.error('[Semantic Cache Error] 寫入語意快取失敗:', err);
}
}
src/app/api/ai/fast-distill/route.ts)現在將 L1 語意快取與 L2 Context Cache 整合至高併發 API 端點中:
// src/app/api/ai/fast-distill/route.ts
import { NextRequest, NextResponse } from 'next/server';
import { searchSemanticCache, saveSemanticCache } from '@/lib/cache/semantic-cache';
import { createOrGetContextCache, generateWithContextCache } from '@/lib/gemini/context-cache';
import { GoogleGenerativeAI } from '@google/generative-ai';
const genAI = new GoogleGenerativeAI(process.env.GEMINI_API_KEY || '');
export async function POST(req: NextRequest) {
const { assetId, fileUri, mimeType, prompt, tokenCount } = await req.json();
// 1. 檢查 L1: Qdrant 語意快取 (目標:< 50ms 回應)
const cachedResponse = await searchSemanticCache(assetId, prompt);
if (cachedResponse) {
return NextResponse.json({
source: 'L1_SEMANTIC_CACHE',
costSaved: '100%',
latency: '<50ms',
result: JSON.parse(cachedResponse),
});
}
let finalResultText = '';
let cacheSource = 'NONE';
// 2. 檢查長脈絡條件 (大於 32,768 Tokens 啟用 L2 Gemini Context Cache)
if (tokenCount > 32768) {
const cacheName = await createOrGetContextCache(assetId, fileUri, mimeType);
if (cacheName) {
console.log('[Pipeline] 觸發 L2 Gemini Context Cache 推論...');
const { text, usageMetadata } = await generateWithContextCache(cacheName, prompt);
finalResultText = text;
cacheSource = 'L2_GEMINI_CONTEXT_CACHE';
console.log(`[Usage Metadata] Context Cache Token Stats:`, usageMetadata);
}
}
// 3. 降級備用方案:一般 Gemini API 呼叫
if (!finalResultText) {
console.log('[Pipeline] 執行標準 Gemini API 呼叫...');
const model = genAI.getGenerativeModel({ model: 'gemini-1.5-flash-002' });
const res = await model.generateContent([
{ fileData: { fileUri, mimeType } },
prompt,
]);
finalResultText = res.response.text();
cacheSource = 'DIRECT_GEMINI_CALL';
}
// 4. 異步更新 L1 語意快取
saveSemanticCache(assetId, prompt, finalResultText).catch(console.error);
return NextResponse.json({
source: cacheSource,
result: JSON.parse(finalResultText),
});
}
我們使用一份包含 150,000 Tokens 的全集 Podcast 逐字稿進行 100 次併發測試,比較三種模式下的效能表現:
| 請求情境 | 回應延遲 (Latency) | 輸入 Token 費用 | 效能與成本評估 |
|---|---|---|---|
| 無快取 (Direct Call) | ~3,800 ms | $0.015 / 次 | 基準線 (100% 費用) |
| L2 Gemini Context Cache | ~850 ms | $0.00375 / 次 | 節省 75% 成本,TTFT 加速 4.4 倍 |
| L1 Qdrant Semantic Cache | ~38 ms | $0.0000 / 次 | 節省 100% 成本,毫秒級秒開 |
今天我們完成了全棧 AI SaaS 上線營運前最 critical 的效能與成本優化工程:
截至今天,我們的 Web 端已具備流暢的 UI、完整的金流、安全的護欄與極速的後端效能!但使用者不可能永遠開著瀏覽器網頁。如果使用者在瀏覽 YouTube、閱讀外文論文或觀看 X (Twitter) 貼文時,想要隨時隨地一鍵提煉,該怎麼辦?
👉 明天(Day 22),我們將進入【跨平台生態擴充篇】:實戰 Chrome Extension 瀏覽器外掛開發與 Side Panel 雙端同步!看我們如何將 OmniVibe AI 的提煉引擎直接注入使用者的瀏覽器中!
我們明天見!🔥