在昨天,我們搞定了圖片傳入到程式中,也儲存到了記憶體中(attached_images 之中),最後把它在 get_input() 中打包並回傳出來。
今天我們要把圖片正式地傳入模型中,接著探討模型支援度的問題,最後,做刪除圖片快捷鍵功能的前半部分。
chat() 要加入 images 參數,然後在有傳入的情況下,要加入 images 鍵:
# src/meowgent/agent.py
class Agent():
...
def chat(..., images: Optional[List[str]] = None) -> ...:
user_msg = {
"role": "user",
"content": user_input
}
if images:
user_msg["images"] = images
self.history_messages.append(user_msg)
原先多輪部分為:
# src/meowgent/agent.py(舊) class Agent(): ... def chat(...) -> ...: self.history_messages.append( { "role": "user", "content": user_input } ) # 使用者輸入加入多輪
main.py 中要接收 images 串列,並把它傳入 chat():
# src/meowgent/cli/main.py
...
if __name__ == "__main__":
...
try:
while True:
user_input, images = get_input() # 接收
...
with ...:
for ... in model.chat(..., images=images):
沒錯,就只要這樣就好了,在 ollama_provider.py 中不用做修改了(因為圖片是以路徑的方式附加在 history_messages 裡的)。
現在,你可以試試看把圖片拖進畫面裡,跟自己的 Agent 來問問題了!
但是,現在還有一個問題,有可能會出現 ollama._types.ResponseError: model "qwen2.5-coder:7b" does not support images (status code: 400) 的報錯,原因是當前的模型不支援視覺,也就是說不是每一個模型都支援圖片多模態能力,所以:
我們對更改模型的指令做修改,讓使用者變換模型時可以得知哪些模型支援圖片,哪些不支援。
原先的程式長這樣:
# src/meowgent/cli/commands.py(舊)
...
@cmd_registry(...)
def _change_model(...) -> ...:
try:
new_model = questionary.select(
"選擇模型",
choices=[m.model for m in ollama.list()["models"]] + ["取消"],
style=...
).ask()
...
現在,要對 choices 動手腳,把支援度加上去,先把原先獲取的串列生成式拉出,改為一般的迴圈。
這裡使用 ollama.show() 之中的 capabilities 屬性獲得模型能力的串列:
特別加上空串列是因為 ollama 在不知道模型能力時,會填上 Pydantic 定義的 None,看到原始碼:
class ShowResponse(SubscriptableBaseModel): ... capabilities: Optional[List[str]] = None # 預設值是 None!
# src/meowgent/cli/commands.py
...
@cmd_registry(...)
def _change_model(...) -> ...:
try:
model_list = []
for m in ollama.list()["models"]:
cap = ollama.show(m.model).capabilities or []
m.model是模型的名稱。
把 cap 拿來判斷是否具有「vision」能力,而因為這邊選項跟選擇後要賦的值不同,所以又用到了 questionary.Choice():
# src/meowgent/cli/commands.py
...
@cmd_registry(...)
def _change_model(...) -> ...:
try:
model_list = []
for ...:
...
model_list.append(questionary.Choice(
title=("[可讀取圖片] " if "vision" in cap else "[不支援圖片] ") + m.model,
value=m.model
))
最後,把 model_list 放入 questionary.select(),也別忘了取消選項,然後順便補上預設:
# src/meowgent/cli/commands.py
...
@cmd_registry(...)
def _change_model(...) -> ...:
try:
new_model = questionary.select(
"選擇模型",
choices=model_list + ["取消"],
default=config.models.default_model, # 順便把預設補上
style=...
).ask()
把預設補上就可以做到提示目前使用模型的效果(上面圖中白底的就是目前所用模型)。
這裡,我們改一下之前的一個小 bug,你可以試試看在選擇的情況下按下 ctrl c,會回傳 None,然後,模型就不能回答了(因為程式不知道要用誰了)。
解決方式就是在原先取消的邏輯加上「new_model 為 None」的情況:
# src/meowgent/cli/commands.py
...
@cmd_registry(...)
def _change_model(...) -> ...:
try:
...
except ...:
...
if new_model in ("取消", None):
return None
這裡的邏輯是「如果
new_model在"取消"或None裡」。
前面說過,目前情況下,若是將圖片傳入沒有視覺能力的模型中,ollama SDK 會拋出 ollama._types.ResponseError: model "qwen2.5-coder:7b" does not support images (status code: 400) 的錯誤,並造成程式崩潰。
現在,我們來處理這件事情。
這裡我不採用例外攔截的方式,而是在傳入圖片時就判斷此模型是否擁有視覺能力,直接攔截不支援的模型:
capabilities後面一樣用空串列兜底。
# src/meowgent/cli/main.py
from ...
import ollama
...
if __name__ == "__main__":
try:
while ...:
...
if not user_input and not images:
...
if images:
if "vision" not in (ollama.show(model.model_name).capabilities or []): # 不支援視覺時
images = None
cli.console.print(cli.render_not_support_vision())
user_input = "[系統提示:使用者原本附帶了圖片,但當前模型不支援視覺讀取,圖片已被移除。請盡可能根據文字問題回答,並適度提醒使用者切換至視覺模型]\n\n" + user_input
render_not_support_vision() 的部分如下:
# src/meowgent/cli/renderers.py
...
class CLIRenderer:
...
def render_not_support_vision(self) -> Padding:
return Padding(
"[red]此模型不支援圖片[/red][dim],若要讀取圖片請用 /model 切換至支援的模型[/dim]",
(0, 0, 0, 2)
)
當我們誤把一張不需要的圖片拖入了畫面中,那要怎麼刪掉呢?
我想要的效果是按下 ctrl + O 後可以開啟刪除圖片的選單,而在選單中將選擇游標移到該圖片選項時,再按一次 ctrl + O 可以打開圖片做預覽。
這邊你也許會想到要用指令的方式,但打字打一打還需要輸入指令顯然不太合理。
首先,我們來做負責選單的函數。
這裡用 async 宣告為「非同步」函數,至於為什麼,我們之後再說(明天為各位解釋),「先把它認定為一個普通函數就好」。
這邊匯入了 attached_images 並建立了 Console 物件(這裡就不從 renderers.py 匯入 cli 了,否則很麻煩):
# src/meowgent/cli/input_prompt.py
from ...
from rich.padding import Padding
from rich.console import Console
from prompt_toolkit.key_binding import KeyBindings
from questionary.prompts.common import InquirerControl
from prompt_toolkit.application.run_in_terminal import in_terminal
import subprocess
...
async def _del_image():
""" 刪除加到對話中的圖片 """
from cli import attached_images
console = Console()
這邊順便先把明天要用的內容一起匯入了。
按下 ctrl + O 後,會有幾個情況:
return。questionary.select() 來做選單了,然後也加上 ctrl + O 預覽的功能。當然選單裡會有「取消」的選項,讓使用者在多張時可以只刪除其一就退出。
在做上面所說的邏輯前,先處理一下渲染刪除提示的部分:
del_img是後面在questionary.select()所選出的項目(但在情況二時則是要自行賦值)。
為了顯示效果,這邊把路徑字串轉為 Path 物件然後只留最後面(檔名 + 副檔名部分)。
# src/meowgent/cli/input_prompt.py
...
async def _del_image():
...
def _render_del_msg():
console.print(Padding(
f"[red]{Path(del_img).name} 已移除[/red]",
(0, 0, 0, 2)
))
這邊,我們把第一及第二點處理好:
第二點的部分,del_img 要先準備好(因為沒有 questionary.select() 了),然後就可以 clear() 把串列清空,最後調用 _render_del_msg() 以及 return 了。
# src/meowgent/cli/input_prompt.py
...
async def _del_image():
...
if not attached_images: # 第一點
console.print(Padding(
"[dim]無圖片[/dim]",
(0, 0, 0, 2)
))
return
elif len(attached_images) == 1:
# 第二點,如果只有一張 -> 直接刪
del_img = attached_images[0]
attached_images.clear()
_render_del_msg()
return
目前我們完成了圖片多模態的功能,已經可以把圖片交到模型手上了,而且也可以辨識及攔截不支援的模型,
但刪除圖片的功能只完成了一小部分,明天繼續完成吧!