iT邦幫忙

2026 iThome 鐵人賽

DAY 22
0
AI Security

打開黑盒子:大型語言模型的機制解釋性入門系列 第 22

Day 22|如果 Concept 的形狀不是線性的,那應該怎麼 Steer?

  • 分享至 

  • xImage
  •  

上一篇我們看到,模型內部的 concept 可能不是一條 direction,也不是一個平坦的 subspace,而是一個具有彎曲結構的 manifold。這對 Activation Steering 會帶來一個直接的問題。標準 steering 通常會計算起點與終點 activation 的差,再沿著同一條 direction 移動:

https://ithelp.ithome.com.tw/upload/images/20260918/20183469H6M25S2soq.png

這條路徑是兩點之間的直線。但如果模型自然形成的 representation 是一個圓、一條曲線,或更複雜的 surface,兩點之間的直線可能會離開模型平常使用的 activation region。也就是說我們雖然把模型推向了正確終點,中間經過的卻可能是模型從未自然形成過的狀態。那麼,與其直接穿過 manifold,我們能不能沿著它原本的形狀移動?

這就是 Manifold Steering 的基本想法。

從 Monday 到 Thursday:走弦,還是走圓弧?

想像模型用一個圓表示星期:

https://ithelp.ithome.com.tw/upload/images/20260918/201834690Qev7wjIZx.png

現在我們想把模型從 Monday steer 到 Thursday。Linear steering 會直接畫一條直線:

Monday ─────────→ Thursday

從幾何上來說,這條線是圓上的一條弦。它確實能從起點抵達終點,卻會穿過圓的內部。如果沿著 weekday manifold 移動,路徑則可能是:

Monday → Tuesday → Wednesday → Thursday

這兩條路最後都會到達 Thursday,但它們對模型 behavior 的影響可能完全不同。Linear path 的中點不一定對應任何自然的 weekday representation;manifold path 上的每個位置,則仍然位於模型平常用來表示星期的區域附近。

先替 Activation Geometry 畫一條路

Manifold Steering 的第一步,是從自然產生的 activations 中估計 representation 的形狀。以 weekdays 為例,我們可以準備很多問題:

What day comes two days after Monday?
What day comes four days after Thursday?
What day comes one day after Saturday?
……

接著依照正確答案,把模型某一層的 activation 分組:

答案是 Monday 的 activations
答案是 Tuesday 的 activations
……
答案是 Sunday 的 activations

每一組取平均後,可以得到七個 activation centroids:

https://ithelp.ithome.com.tw/upload/images/20260918/20183469E2aq4ikjML.png

如果這些 centroids 形成循環結構,研究者就可以在它們之間 fitting 一條平滑的 periodic spline,近似模型的 weekday manifold。[1] 可以把這條曲線寫成:

https://ithelp.ithome.com.tw/upload/images/20260918/20183469OvYVwXQqJE.png

其中 u 是 manifold 自己的 intrinsic coordinate。Linear steering 直接在原本的 activation space 中插值:

https://ithelp.ithome.com.tw/upload/images/20260918/20183469lBRVznJGnM.png

Manifold steering 則先在 intrinsic coordinate 中移動,再把每個位置 decode 回 activation space:

https://ithelp.ithome.com.tw/upload/images/20260918/20183469QLgl3cRjfY.png

核心差異只有:

Linear steering → 直接在高維 activation space 中畫直線

Manifold steering → 先沿著 concept 自己的座標移動,再回到 activation space

把 Path 上的每個點塞回模型

有了兩條路徑後,研究者可以真的做 intervention。假設我們把 Monday 到 Thursday 的路徑切成 50 個 waypoints。對每一個 waypoint:

  1. 執行一個正常 prompt
  2. 跑到指定 layer 與 token position
  3. 把 Residual Stream 替換成該 waypoint 的 activation
  4. 繼續完成 forward pass
  5. 記錄模型對七個 weekday tokens 的 probability

簡化成 pseudocode,大致是:

for t in path_positions:
    hidden = steering_path(t)

    logits = continue_forward_with_patch(
        prompt,
        layer=target_layer,
        position=-1,
        replacement=hidden,
    )

    weekday_probs.append(
        get_concept_probabilities(logits)
    )

這和之前的 Activation Patching 非常相似。不同的是,過去我們通常只比較 clean 與 corrupted 兩個 activation;現在則沿著一整條路徑,連續觀察模型 behavior 如何改變。

Linear Steering 可能會「瞬間傳送」

Goodfire 在 weekdays、months、letters 與 ages 等結構化任務中比較了 linear steering 和 manifold steering。[1] 以星期為例,manifold steering 產生的 output probability 會較平順地沿著相鄰概念移動:

Monday
↓
Tuesday
↓
Wednesday
↓
Thursday

Linear steering 則可能出現比較不自然的「teleportation」:

Monday
↓
Wednesday
↓
某個不相干 token
↓
Thursday

也就是 probability mass 並沒有按照星期的自然順序移動,而是在不同概念之間跳躍,甚至暫時落到和任務無關的 tokens 上。這個結果提供了一個很直覺的解釋:

Linear steering 抵達了終點,卻可能穿過低密度、off-manifold 的內部狀態。

而 manifold steering 比較像沿著模型自然形成的 representation 移動,因此產生的中間 behaviors 也更接近模型在沒有 intervention 時會自然形成的輸出。[1]

同一個 Concept,其實有兩張地圖

這項研究還發現了一個更有趣的現象。

我們剛才討論的是 activation space:模型內部的 hidden representations 如何排列。

但對每個 prompt,模型也會產生一個 output probability distribution。例如,當答案是 Tuesday 時,模型可能:

Tuesday     0.78
Monday      0.08
Wednesday   0.07
其他 tokens 0.07

當答案是 Wednesday 時:

Wednesday   0.81
Tuesday     0.06
Thursday    0.06
其他 tokens 0.07

如果把每一個 probability distribution 也視為一個點,這些 points 同樣會形成一個 behavior manifold。在 weekday task 中:

  • activation manifold 呈現圓形;
  • output distributions 形成的 behavior manifold 也呈現圓形;
  • 兩者沿 manifold 測量出的距離彼此高度對應。[1]

這代表模型內部 representation 的 geometry,不一定只是漂亮的 visualization。它可能真的約束了模型能夠產生哪些 behaviors,以及 behaviors 之間如何自然地轉換。

Primal Space 與 Dual Space

這裡也可以用另外一種方式理解 steering。

Primal:從 Activation 本身出發

我們可以直接觀察自然 activations 分布在哪裡?接著 fitting activation manifold,再沿著它移動。這可以稱為 activation-aware steering,或 primal-space 的觀點。

Dual:從 Readout 或 Output 出發

另一種做法不是直接研究 activation centroids,而是研究:哪些 directions 可以讀出不同的 concept values?這些 readout directions 彼此形成什麼 geometry?我們希望模型最後形成什麼 output distribution?這可以稱為 field-aware 或 dual-space 的觀點。[2]

例如,我們可以替不同目標值各自訓練 probe:

Probe for 300
Probe for 400
Probe for 500
……

再觀察這些 probe directions 彼此之間的關係,建立一個比較平滑的 readout field。或者更直接地,先在 behavior space 中指定一條理想路徑:

Monday distribution
↓
Tuesday distribution
↓
Wednesday distribution
↓
Thursday distribution

再反向尋找什麼樣的 activation path,能讓模型產生這條 behavior trajectory?

Goodfire 將這種從 behavior geometry 反推 activation path 的方法稱為 pullback。得到的 activation path,會和直接從自然 activations fitting 出來的 manifold 相似。[1]這提供了兩個互相呼應的視角:

Representation → Behavior
沿 activation manifold steer,看看輸出怎麼變

Behavior → Representation
先指定自然 output path,再反推 activation path

這裡的「dual」指的是從 readout/distribution 的角度研究同一個 concept,不需要把它理解成嚴格線性代數中唯一固定的 dual space。

Geometry 不只讓 Steering 更自然,也提供了因果證據

上一篇我們找到 weekday circle 時,仍然可以懷疑這可能只是 PCA 畫出來的漂亮形狀。但 Manifold Steering 真正有價值的地方,是它把 geometry 變成可以 intervention 的 hypothesis。如果 activation manifold 只是偶然的視覺 pattern,那麼沿著它走,不應該特別產生自然、有順序的 behavior。

但實驗中:

  • activation geometry 與 behavior geometry 呈現相似結構;
  • 沿 activation manifold steering,output 也沿 behavior manifold 移動;
  • 先指定 behavior path 反推 activation,則又會找回近似的 activation geometry。[1]

因此,geometry 不只是在描述 representation 長什麼樣。它可能是模型實際組織並控制 behavior 的因果結構。

這也把 steering 問題重新定義了。過去我們問的是我要找哪一條 direction?現在更一般的問題變成我要找的是哪一種 geometry?

Linear Steering 並沒有因此失效

Manifold Steering 並不表示所有 linear steering 都是錯的。任何平滑 manifold 在足夠小的局部範圍內,都可以近似成平面。因此,當:

  • intervention 很小;
  • 起點與終點很接近;
  • concept geometry 本來就接近線性;
  • 我們只在單一 context 附近操作;

一條 local direction 可能已經非常有效。Linear steering 的優點仍然很明顯:

  • 計算便宜
  • 不需要先 fitting manifold
  • 容易套用到大量 prompts
  • 不需要已知 concept topology
  • implementation 簡單

Manifold steering 則需要:

  • 足夠的自然 activation samples
  • 選擇 layer 與 token position
  • 找到 concept 的 intrinsic coordinates
  • fitting 一個穩定的 manifold
  • 設計合理的 interpolation path

所以它不是 linear steering 的全面替代品。比較合理的說法是:

Linear steering 是一個方便的近似;當 representation curvature 變得重要時,geometry-aware steering 才可能帶來額外價值。

我們還不知道真實世界的 Concept 長什麼樣

目前 Manifold Steering 最清楚的結果,多半來自具有已知結構的 tasks:

  • weekdays
  • months
  • letters
  • ages
  • grid 或 graph environments

這些任務有明確的順序、循環或座標,研究者知道該怎麼替 manifold 定義 intrinsic coordinates。但像:

  • honesty
  • harmfulness
  • persona
  • political belief
  • deception
  • emotional attachment

可能沒有單一、乾淨的 coordinate system。它們可能:

  • 隨 context 改變
  • 有多條分支
  • 由多個 manifolds 組成
  • 在不同 layers 使用不同 geometry
  • 根本無法由一條低維 smooth surface 充分描述

此外,一條「更接近自然 output distribution」的 path,也不一定就是更安全、更正確的 path。因此,Manifold Steering 目前最好被理解成:

一個證明 representation geometry 值得被認真對待的研究方向,而不是已經成熟的通用控制工具。

2026 年也陸續出現其他 geometry-aware steering 方法,例如利用 kernel space 尋找非線性路徑的 Curveball Steering、保留 activation norm 的 Spherical Steering,以及把 steering 寫成 Riemannian geodesic optimization 的方法。[3][4] 這個領域仍然發展得非常快。

一堆方法分析起來好麻煩,我們不能直接看模型自己怎麼解釋嗎?

過去幾天,我們已經看到很多解讀、控制模型的方法,它們可以閱讀、改變模型的 internal state,進而改變最後的 behavior。但一路上我們都一直強調解讀需要的嚴謹性,以及每個方法給我們的證據的有限性,或許你會想:「我不能直接問模型他為什麼這樣回答嗎?」或是「我不能直接看模型的 CoT 嗎?」。

但不管是哪個方法,我們該怎麼知道模型說的是不是實話?甚至,他自己知不知道自己說的是不是實話?他到底有沒有能力得知自己內部的資訊?假設 steering 讓模型的回答從 Yes 變成 No,接著模型又寫出一段看似完整的理由,解釋自己為什麼選擇 No,這段理由真的反映了造成答案改變的內部 computation 嗎?還是模型只是在答案決定之後,生成一段聽起來合理的文字?

下一篇,我們會開始研究模型自己提供的 explanation:Chain-of-Thought,真的是模型的推理過程嗎?

參考資料與延伸閱讀

[1] Wurgaft, D., Rager, C., Kowal, M., et al., “Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior”, 2026. https://arxiv.org/abs/2605.05115

[2] Sarfati, R., et al., “The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models’ Posteriors”, 2026. https://arxiv.org/abs/2602.02315

[3] Raval, S., Song, H. J., Wu, L., et al., “Curveball Steering: The Right Direction To Steer Isn’t Always Linear”, 2026. https://arxiv.org/abs/2603.09313

[4] You, Z., Deng, C., & Chen, H., “Spherical Steering: Geometry-Aware Activation Rotation for Language Models”, 2026. https://arxiv.org/abs/2602.08169


上一篇
# Day 21|模型內的 Concept 到底是什麼形狀?
下一篇
Day 23|模型的 CoT,可以解釋他的行為嗎?
系列文
打開黑盒子:大型語言模型的機制解釋性入門27
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言