在傳統的線上學習情境中,學生在學習數學或理化計算題卡關時,往往面臨一個最大的痛點:「很難把複雜的方程式或自己寫到一半的計算過程完整描述給 AI 聽」。如果只能在對話框中敲出「3x+5=20,我下一步該怎麼移項?」特殊符號也很難手動輸入,不僅輸入效率低落,AI 也無法真正「看見」學生的筆跡邏輯與計算盲點。
因此,在今天的實作中,我們將結合 Gemini Vision API 與前端的檔案上傳介面,並透過 Postman 進行 API 介面測試,讓我們的 AI 虛擬助教具備「眼睛」,能夠直接看懂學生上傳的手寫算式照片。
multer 處理檔案上傳,並透過 Google GenAI SDK 與 Gemini 模型串接圖片解析。在專案根目錄執行以下指令,安裝用於處理 multipart/form-data 檔案上傳的 multer:
npm install multer
services/visionService.js)我們將多模態的處理邏輯獨立模組化,利用 Gemini 的圖片解析能力:
import { GoogleGenAI } from '@google/genai';
// 初始化 Google GenAI 用戶端(會自動讀取 process.env.GEMINI_API_KEY)
const ai = new GoogleGenAI();
/**
* 處理學生上傳的手寫算式圖片,進行引導式批改(含自動重試機制防範 503 忙線)
* @param {Object} file - Express-fileupload 或 multer 傳入的檔案物件
* @param {number} retries - 重試次數,預設 3 次
* @param {number} delay - 每次重試間隔時間 (毫秒),預設 2000ms
* @returns {Promise<string>} AI 的引導回饋
*/
export async function analyzeHandwriting(file, retries = 3, delay = 2000) {
// 將上傳的 Buffer 轉換成 Gemini API 支援的 inlineData 格式
const imagePart = {
inlineData: {
data: file.buffer.toString("base64"),
mimeType: file.mimetype,
},
};
const prompt = (
"你是一個具備同理心的 AI 虛擬助教。請看這張學生的手寫算式或作業圖片," +
"找出學生觀念錯誤的地方。**請注意:絕對不能直接給出正確答案**," +
"而是要用蘇格拉底式的提問法,引導學生自己發現錯誤並思考下一步。" +
"重要排版限制:請全部使用純文字與一般標點符號回答,絕對不要使用 LaTeX 語法、絕對不能出現 $ 符號(例如數學式請直接寫 $x$ 改寫為 x)。"
);
for (let i = 0; i < retries; i++) {
try {
// 呼叫支援多模態的 gemini-1.5-flash 模型
const response = await ai.models.generateContent({
model: 'gemini-3.6-flash',
contents: [prompt, imagePart],
});
return response.text;
} catch (error) {
console.warn(`⚠️ 嘗試第 ${i + 1} 次呼叫 Gemini Vision 失敗:`, error.message);
// 若遇到 503 伺服器忙線,且還沒達到最大重試次數,則等待後繼續
if (error.status === 503 && i < retries - 1) {
console.log(`⏳ 伺服器忙線中,將在 ${delay / 1000} 秒後進行第 ${i + 2} 次重試...`);
await new Promise((resolve) => setTimeout(resolve, delay));
continue;
}
// 若不是 503 或重試用盡,則直接拋出錯誤
throw error;
}
}
}
routes/visionRoutes.js)設定 API 端點來接收網頁或 Postman 傳遞的圖片檔案:
import express from 'express';
import multer from 'multer';
import { analyzeHandwriting } from '../services/visionService.js';
const router = express.Router();
// 使用記憶體儲存機制 (Memory Storage) 暫存上傳的圖片
const upload = multer({ storage: multer.memoryStorage() });
// POST /api/vision/tutoring
router.post('/tutoring', upload.single('image'), async (req, res, next) => {
try {
if (!req.file) {
return res.status(400).json({ error: '沒有上傳圖片檔案' });
}
const feedback = await analyzeHandwriting(req.file);
res.json({
success: true,
ai_feedback: feedback
});
} catch (error) {
next(error); // 交由 Express 全域錯誤處理中間件處理
}
});
export default router;
為了確保後端 Vision API 能正確運作,在對接前端介面前,我們使用 Postman 進行獨立驗證:
POST,網址輸入 http://localhost:3000/api/vision/tutoring。form-data 頁籤,欄位名稱輸入 image,並將右側的文字下拉選單從 Text 改為 File。
當後端透過 Postman 驗證成功後,我們即可實作前端介面。透過隱藏式 <input type="file"> 綁定按鈕,並利用 FileReader 讓學生在上傳後能立即在聊天室看見預覽卡片,再透過 FormData 將圖檔送往後端:
// 處理手寫算式圖片上傳與 Gemini Vision 串接
const imageInput = document.getElementById('imageInput');
if (imageInput) {
imageInput.addEventListener('change', async (event) => {
const file = event.target.files[0];
if (!file || isGenerating) return;
// A. 在聊天室畫面上即時顯示學生上傳的照片預覽
const reader = new FileReader();
reader.onload = function(e){
const userPreviewHtml = `
<div class="flex gap-3 md:gap-4 max-w-[92%] md:max-w-2xl self-end flex-row-reverse animate-fade-in">
<div class="flex flex-col gap-1.5 items-end">
<div class="px-1 text-xs text-warmbg-muted font-semibold">你 - 上傳手寫算式</div>
<div class="bg-brand-500 text-white rounded-3xl rounded-tr-lg p-3 md:p-4">
<img src="${e.target.result}" alt="手寫題目照片" class="rounded-xl max-w-xs max-h-60 object-contain border border-white/20">
</div>
</div>
</div>`;
chatContainer.insertAdjacentHTML('beforeend', userPreviewHtml);
scrollToBottom();
};
reader.readAsDataURL(file);
// B. 顯示思考動畫並透過 FormData 呼叫後端 Vision API
setInteractionState(true);
chatContainer.insertAdjacentHTML('beforeend', loadingBubbleHtml());
scrollToBottom();
const formData = new FormData();
formData.append('image', file);
try {
const response = await fetch('http://localhost:3000/api/vision/tutoring', {
method: 'POST',
body: formData
});
if (!response.ok) throw new Error(`HTTP error! status: ${response.status}`);
const result = await response.json();
const data = result.data || result;
document.getElementById('loading-bubble')?.remove();
// C. 顯示 AI 助教的視覺引導回饋
const feedbackText = data.ai_feedback || data.reply || '老師收到你的手寫算式了,我們來一起看看這題!';
chatContainer.insertAdjacentHTML('beforeend', aiBubbleHtml(feedbackText, [], 'breakdown'));
scrollToBottom();
} catch (error) {
console.error('Vision API error:', error);
document.getElementById('loading-bubble')?.remove();
chatContainer.insertAdjacentHTML('beforeend', aiBubbleHtml('抱歉,圖片辨識失敗了,請重新上傳!', []));
scrollToBottom();
} finally {
setInteractionState(false);
imageInput.value = '';
}
});
}
\n 或 $ 符號,看起來像排版錯亂。\n 呈現。只要未來在前端網頁透過專門的解析工具(如 Marked.js 或 MathJax)渲染,這些符號就會自動轉化為漂亮的段落與數學公式!現在有了多模態,學生不需要花大量時間把複雜的方程式或手寫計算過程打成文字,只要隨手拍下一張筆記或考卷,AI 就能夠精準「看懂」並找出學生的思考卡點,降低了科技輔助學習的門檻,也讓 AI 助教更貼近真實課堂中,就像老師親切看著學生算草稿、從旁給予關鍵提示一樣。
明天我們將進入基礎 RAG(Retrieval-Augmented Generation) 的概念,讓 AI 助教能夠讀取指定的 PDF 教材,避免 AI 在面對艱澀專業科目時「瞎編答案」。