前幾天 BehaviorGuard 已經完成了 Process、Command、Network、File 與 Persistence Detection。
目前系統已經能針對不同類型的端點行為產生 Alert。
| Detection 類型 | 目前功能 |
|---|---|
| Process Detection | 偵測可疑 Process 與 Parent / Child Process 關係 |
| Command Detection | 偵測 curl | bash、python3 -c 等可疑指令 |
| Network Detection | 偵測 Listener、External Connection、Interpreter Network Activity |
| File Detection | 偵測敏感檔案修改、刪除、隱藏檔案與權限變化 |
| Persistence Detection | 偵測 Cron、systemd、Shell Profile、SSH authorized_keys 等持久化行為 |
但是做到這裡後,我開始遇到一個問題。
假設 BehaviorGuard 同時產生:
BG-CMD-002
Python Inline Code Execution
以及:
BG-NET-003
Unexpected External Connection
我們現在知道「有兩個可疑 Alert」。
但是下一個問題是:
這兩個 Alert 是兩件完全沒有關係的事情,還是其實屬於同一次攻擊?
因此 Day14 開始進入 BehaviorGuard 很重要的一個功能:
前面的 BehaviorGuard 流程比較像:
Telemetry
↓
Detection Rule
↓
Alert
例如 Command Detection 發現:
python3 -c "..."
↓
BG-CMD-002
Python Inline Code Execution
↓
MEDIUM
Network Detection 又發現:
External Connection
↓
BG-NET-003
Unexpected External Connection
↓
MEDIUM
如果把兩個事件分開來看,其實都不一定代表正在遭受攻擊。
例如:
python3 -c "print('hello')"
這可能只是使用者正常測試 Python。
而程式建立 External Connection 也不一定是攻擊,因為很多正常軟體本來就需要連線到 Internet。
但如果今天看到的是:
python3 -c
↓
幾秒後
↓
Unexpected External Connection
而且兩個 Alert 都來自:
同一個 PID
那風險就明顯不同。
因此 Correlation 的核心目的就是:
把多個獨立 Alert 串起來,判斷它們是否可能屬於同一條攻擊行為。
做到 Correlation 之後,我開始更清楚了解三個常出現在 SOC 與 EDR 裡面的名詞。
| 名稱 | 意思 | BehaviorGuard 範例 |
|---|---|---|
| Event | 系統發生的一件事情 | Process 啟動、檔案修改、網路連線 |
| Alert | Detection Rule 判斷 Event 可疑 | python3 -c 觸發 BG-CMD-002 |
| Incident | 多個相關 Alert 組成的安全事件 | 可疑 Python 執行後又建立外部連線 |
整體流程可以理解成:
Event
↓
Detection
↓
Alert
↓
Correlation
↓
Incident
這也讓我理解,真實 SOC 並不是看到一個 Alert 就馬上認定系統遭受攻擊。
因為單一 Alert 很可能只是 False Positive。
更重要的是:
多個 Alert 之間有沒有關係?
要做事件關聯,第一個問題就是:
BehaviorGuard 要怎麼知道「剛剛發生過什麼」?
因此我先建立:
alert_buffer = []
它的作用就是暫時保存最近收到的 Detection Alert。
例如目前可能保存:
Alert Buffer
BG-CMD-002
BG-NET-005
BG-NET-003
當新的 Alert 進來時,會呼叫:
def add_alert(alert):
alert["correlation_time"] = time.time()
alert_buffer.append(alert)
cleanup_old_alerts()
check_correlation_rules()
這段程式可以拆成四個步驟。
| 程式 | 功能 |
|---|---|
time.time() |
記錄 Alert 進入 Correlation Engine 的時間 |
alert_buffer.append(alert) |
把 Alert 放進暫存區 |
cleanup_old_alerts() |
清除太舊的 Alert |
check_correlation_rules() |
開始檢查事件之間是否符合 Correlation Rule |
整個流程:
New Alert
↓
記錄時間
↓
Alert Buffer
↓
清除舊事件
↓
Correlation Rule
Alert Buffer 可以把它想成 BehaviorGuard 的「短期記憶」。
如果系統完全不保留前面的 Alert,那新的 Alert 出現時,就沒有辦法知道前面幾秒鐘曾經發生過什麼事情。
接著遇到第二個問題。
Alert 不可能永久存在於 Buffer。
假設:
09:00
BG-CMD-002
到了:
20:00
BG-NET-003
如果因為兩個 Rule 看起來有關,就認為它們是同一次攻擊,False Positive 一定會非常高。
因此我設定:
TIME_WINDOW = 30
代表目前 BehaviorGuard 的第一版 Correlation,只會考慮:
30 秒內
發生的事件。
例如:
19:20:01 BG-CMD-002
19:20:04 BG-NET-003
兩個事件相隔:
3 秒
因此有可能存在關聯。
但是:
19:20:01 BG-CMD-002
20:30:00 BG-NET-003
相隔一個多小時,就不應該直接被串在一起。
清除舊 Alert 的程式:
def cleanup_old_alerts():
current_time = time.time()
global alert_buffer
alert_buffer = [
alert
for alert in alert_buffer
if current_time - alert["correlation_time"] <= TIME_WINDOW
]
這段程式實際上就是計算:
現在時間 - Alert 發生時間
如果結果:
<= 30 秒
就留下。
如果:
> 30 秒
就從 Buffer 移除。
這就是:
Time-based Correlation
有了 Time Window 之後,下一個問題是:
我怎麼知道兩個 Alert 是不是同一個程式產生的?
因此第一版我先使用 Linux 的 PID。
PID 是 Process ID,也就是每個 Process 的識別碼。
例如:
python3
PID = 4210
Command Detection 發現:
BG-CMD-002
PID = 4210
接著 Network Detection 又發現:
BG-NET-003
PID = 4210
兩邊 PID 都是:
4210
因此有很強的理由認為:
這兩個行為是同一個 Process 做的。
第一版的條件就是:
30 秒內
+
Same PID
測試程式:
alert1 = {
"rule": "BG-CMD-002",
"name": "Python Inline Code Execution",
"severity": "MEDIUM",
"pid": 4210
}
alert2 = {
"rule": "BG-NET-003",
"name": "Unexpected External Connection",
"severity": "MEDIUM",
"pid": 4210
}
送進 Correlation Engine:
add_alert(alert1)
time.sleep(2)
add_alert(alert2)
測試結果:
Sending Alert 1...
Sending Alert 2...
========== BehaviorGuard Correlation ==========
PID : 4210
Current Rule : BG-NET-003
Related Alerts : 1
Alert Chain:
-> BG-CMD-002 (Python Inline Code Execution)
-> BG-NET-003 (Unexpected External Connection)
================================================
到這裡代表:
Time Window ✅
Same PID ✅
第一版事件關聯成功。
第一版雖然成功,但是很快又發現一個問題。
如果邏輯只有:
Same PID
+
30 秒
那代表:
任何 Alert A
+
任何 Alert B
只要 PID 一樣,都會被 BehaviorGuard 認為有關係。
這樣其實太寬鬆。
例如:
BG-NET-001
+
BG-FILE-001
不一定有明確的攻擊邏輯。
因此不能只說:
PID 一樣 = Attack Chain
而是還要加入:
因此新增:
correlation/rules.py
目前 Day14 建立五條 Correlation Rule。
| Rule | Attack Pattern | Severity |
|---|---|---|
| BG-CORR-001 | BG-CMD-001 → BG-NET-003 | HIGH |
| BG-CORR-002 | BG-CMD-002 → BG-NET-003 | HIGH |
| BG-CORR-003 | BG-PROC-001 → BG-NET-005 | CRITICAL |
| BG-CORR-004 | BG-PROC-002 → BG-NET-003 | HIGH |
| BG-CORR-005 | BG-CMD-002 → BG-NET-005 → BG-NET-003 | CRITICAL |
程式:
CORRELATION_RULES = [
{
"id": "BG-CORR-001",
"name": "Download and External Connection",
"severity": "HIGH",
"required_rules": [
"BG-CMD-001",
"BG-NET-003"
],
"description": (
"A suspicious download-to-shell command was followed "
"by an unexpected external network connection."
)
},
{
"id": "BG-CORR-002",
"name": "Python Execution with External Connection",
"severity": "HIGH",
"required_rules": [
"BG-CMD-002",
"BG-NET-003"
],
"description": (
"Python inline code execution was associated with "
"an unexpected external network connection."
)
},
{
"id": "BG-CORR-003",
"name": "Web Shell with Network Activity",
"severity": "CRITICAL",
"required_rules": [
"BG-PROC-001",
"BG-NET-005"
],
"description": (
"A web server spawned a shell which was followed "
"by suspicious shell or interpreter network activity."
)
},
{
"id": "BG-CORR-004",
"name": "Shell Network Tool Chain",
"severity": "HIGH",
"required_rules": [
"BG-PROC-002",
"BG-NET-003"
],
"description": (
"A shell spawned a network-related tool which then "
"established an unexpected external connection."
)
},
{
"id": "BG-CORR-005",
"name": "Suspicious Execution and Network Activity",
"severity": "CRITICAL",
"required_rules": [
"BG-CMD-002",
"BG-NET-005",
"BG-NET-003"
],
"description": (
"Suspicious interpreter execution was followed by "
"network activity and an unexpected external connection."
)
}
]
現在 BehaviorGuard 已經不是:
只要兩個 Alert PID 一樣就關聯
而是:
PID 相同
+
Time Window 內
+
符合特定 Rule Combination
才會產生 Correlated Alert。
第一條實際測試成功的規則:
BG-CORR-002
Python Execution with External Connection
需要:
BG-CMD-002
↓
BG-NET-003
意思就是:
Python Inline Code Execution
↓
Unexpected External Connection
測試結果:
========== BehaviorGuard Correlated Alert ==========
Rule : BG-CORR-002
Name : Python Execution with External Connection
Severity : HIGH
PID : 4210
Description : Python inline code execution was associated with an unexpected external network connection.
Attack Chain:
-> BG-CMD-002 (Python Inline Code Execution)
-> BG-NET-003 (Unexpected External Connection)
=====================================================
這時候 BehaviorGuard 已經不只是說:
我看到兩個 Alert
而是開始說:
我看到 Python 可疑執行之後,
同一個 Process 又出現 External Connection。
這就是 Correlation Rule 比單一 Detection Rule 更有價值的地方。
做到這裡之後,又遇到新的問題。
假設 Correlation Rule 是:
BG-CMD-002
↓
BG-NET-003
如果程式只是檢查:
Buffer 有 BG-CMD-002
Buffer 有 BG-NET-003
那這種順序:
BG-NET-003
↓
BG-CMD-002
也可能被判斷成功。
但是:
Command → Network
跟:
Network → Command
代表的行為並不完全一樣。
因此 Correlation 不能只檢查:
「哪些 Rule 出現過?」
還要判斷:
「Rule 出現的順序對不對?」
因此加入:
def find_rule_sequence(alerts, required_rules):
matched_alerts = []
rule_index = 0
for alert in alerts:
if alert.get("rule") == required_rules[rule_index]:
matched_alerts.append(alert)
rule_index += 1
if rule_index == len(required_rules):
return matched_alerts
return None
這段程式會按照 Alert 的發生時間依序尋找:
第一個 Required Rule
↓
第二個 Required Rule
↓
第三個 Required Rule
只有順序全部正確才算 Match。
正確順序:
BG-CMD-002
↓
BG-NET-003
結果:
BG-CORR-002
成功觸發。
但我另外做了一個反向測試:
Sending Network Alert first...
Sending Command Alert second...
也就是:
BG-NET-003
↓
BG-CMD-002
最後沒有出現:
BG-CORR-002
代表:
Sequence Check ✅
成功。
接下來又遇到一個問題。
假設:
BG-CMD-002
↓
BG-NET-003
已經成功觸發:
BG-CORR-002
但是因為這些 Alert 還存在 Buffer 裡面,如果又有一個新的 Network Alert 進來:
BG-NET-003
Correlation Engine 又重新檢查一次。
它可能再次看到:
BG-CMD-002
+
BG-NET-003
然後再次產生:
BG-CORR-002
最後變成:
BG-CORR-002
BG-CORR-002
BG-CORR-002
BG-CORR-002
這就是:
Duplicate Alert
如果資安系統一直產生相同 Alert:
Alert
Alert
Alert
Alert
Alert
Alert
SOC 分析人員最後可能會開始忽略 Alert。
這種狀況叫:
也就是「告警疲勞」。
因此資安產品並不是:
Alert 越多越好
而是:
Alert 要有意義
因此加入:
triggered_correlations = {}
用來記錄最近已經觸發過的 Correlation。
例如:
BG-CORR-002
PID 6000
會記成類似:
("BG-CORR-002", 6000)
程式:
correlation_key = (
correlation_rule["id"],
pid
)
接著檢查:
if correlation_key in triggered_correlations:
continue
意思就是:
如果這個 PID 的這條 Correlation Rule 最近已經報過,就不要再重複報。
成功觸發後:
triggered_correlations[correlation_key] = time.time()
記錄觸發時間。
測試流程:
BG-CMD-002
↓
BG-NET-003
↓
BG-CORR-002
接著再送一次:
BG-NET-003
測試結果:
Sending Command Alert...
Sending Network Alert...
========== BehaviorGuard Correlated Alert ==========
Rule : BG-CORR-002
...
=====================================================
Sending another Network Alert...
最後:
沒有第二個 BG-CORR-002
代表:
Duplicate Suppression ✅
成功。
完成兩個 Alert 的 Correlation 後,我開始嘗試更完整的 Attack Chain。
這次測試:
Stage 1
BG-CMD-002
Python Inline Code Execution
↓
Stage 2
BG-NET-005
Shell or Interpreter Network Activity
↓
Stage 3
BG-NET-003
Unexpected External Connection
也就是:
Python Inline Execution
↓
Interpreter Network Activity
↓
External Connection
因此建立:
BG-CORR-005
Suspicious Execution and Network Activity
Severity = CRITICAL
測試程式如下:
import time
from correlation.engine import add_alert
alert1 = {
"rule": "BG-CMD-002",
"name": "Python Inline Code Execution",
"severity": "MEDIUM",
"pid": 7000
}
alert2 = {
"rule": "BG-NET-005",
"name": "Shell or Interpreter Network Activity",
"severity": "HIGH",
"pid": 7000
}
alert3 = {
"rule": "BG-NET-003",
"name": "Unexpected External Connection",
"severity": "MEDIUM",
"pid": 7000
}
print("Stage 1: Suspicious Python Execution")
add_alert(alert1)
time.sleep(1)
print("Stage 2: Interpreter Network Activity")
add_alert(alert2)
time.sleep(1)
print("Stage 3: Unexpected External Connection")
add_alert(alert3)
三個 Alert:
PID = 7000
而且按照順序:
BG-CMD-002
↓
BG-NET-005
↓
BG-NET-003
進入 Correlation Engine。
這次 BehaviorGuard 成功產生兩個 Correlated Alert。
第一個:
========== BehaviorGuard Correlated Alert ==========
Rule : BG-CORR-002
Name : Python Execution with External Connection
Severity : HIGH
PID : 7000
Attack Chain:
-> BG-CMD-002 (Python Inline Code Execution)
-> BG-NET-003 (Unexpected External Connection)
=====================================================
第二個:
========== BehaviorGuard Correlated Alert ==========
Rule : BG-CORR-005
Name : Suspicious Execution and Network Activity
Severity : CRITICAL
PID : 7000
Attack Chain:
-> BG-CMD-002 (Python Inline Code Execution)
-> BG-NET-005 (Shell or Interpreter Network Activity)
-> BG-NET-003 (Unexpected External Connection)
=====================================================
這代表:
兩階段行為
BG-CMD-002
↓
BG-NET-003
= HIGH
但當更多證據出現:
BG-CMD-002
↓
BG-NET-005
↓
BG-NET-003
= CRITICAL
BehaviorGuard 可以把風險往上提升。
這次雖然還沒有正式建立 Risk Scoring Engine,但是已經開始出現:
Risk Escalation
也就是「風險升級」。
可以整理成:
| 行為 | Severity |
|---|---|
| 單一 Python Inline Execution | MEDIUM |
| Python Execution + External Connection | HIGH |
| Execution + Interpreter Network + External Connection | CRITICAL |
概念:
單一可疑證據
↓
MEDIUM
更多相關證據
↓
HIGH
完整攻擊行為鏈
↓
CRITICAL
這跟前面直接使用單一 Rule 判斷風險相比,更接近真正 EDR 的思考方式。
因為:
一個行為可疑
不代表一定要立刻 Response。
但是:
可疑 Execution
+
可疑 Network
+
External Communication
同時發生時,可信度就會提高很多。
目前 Day14 的 Correlation Engine 可以整理成:
New Alert
↓
Alert Buffer
↓
30 Seconds Time Window
↓
Group by PID
↓
Correlation Rules
↓
Rule Sequence
↓
Duplicate Check
↓
Correlated Alert
↓
Attack Chain
整個 BehaviorGuard 的架構也慢慢變成:
Linux Endpoint
↓
Telemetry
↓
Detection Engine
↓
Individual Alert
↓
Correlation Engine
↓
Attack Chain
↓
Incident
↓
Risk Scoring
↓
Response
前面的 BehaviorGuard 主要是在回答:
「發生了什麼事情?」
現在開始可以進一步回答:
「這些事情是不是其實互相有關?」
今天實作 Correlation 過程中其實遇到不少問題。
| 問題 | 原因 | 解決方式 |
|---|---|---|
| Python SyntaxError | 不小心寫成 kimport time |
修正成 import time |
| 相同 PID 就全部關聯 | Correlation 條件太寬 | 建立 Correlation Rules |
| 事件反過來也會 Match | 只判斷 Rule 是否存在 | 加入 Rule Sequence |
| Correlation 一直重複出現 | Buffer 仍保留舊 Alert | 加入 Duplicate Suppression |
| BG-CORR-005 一開始沒有出現 | rules.py 尚未加入完整 Rule |
補上 BG-CORR-005 |
我覺得今天最有價值的地方,不只是最後成功印出:
BG-CORR-005
而是實際做之後才發現:
Correlation
不是單純:
Alert A + Alert B
而已。
真正需要考慮:
Time
PID
Rule Combination
Sequence
Duplicate
Severity
雖然 Day14 已經能做多階段 Attack Chain,但是目前還有一個很大的限制。
現在的 Correlation 主要依賴:
Same PID
例如:
BG-CMD-002
PID 7000
BG-NET-005
PID 7000
BG-NET-003
PID 7000
這種情況很好處理。
但是實際 Linux 攻擊可能是:
bash
PID 7000
↓
python3
PID 7010
↓
nc
PID 7020
三個 PID 都不一樣。
但是它們可能存在:
Parent
↓
Child
↓
Grandchild
的關係。
也就是:
Process Tree
人類可以看出:
bash
↓
python3
↓
nc
很可能屬於同一條行為鏈。
但目前的 BehaviorGuard 因為:
PID 7000
PID 7010
PID 7020
不一樣,所以還無法完整關聯。
| 功能 | 狀態 |
|---|---|
| Alert Buffer | ✅ |
| Time Window | ✅ |
| Same PID Correlation | ✅ |
| Correlation Rules | ✅ |
| Rule Sequence | ✅ |
| Duplicate Suppression | ✅ |
| Two-stage Correlation | ✅ |
| Three-stage Correlation | ✅ |
| Attack Chain Output | ✅ |
| Risk Escalation Concept | ✅ |
Detection 解決:
「這個行為可不可疑?」
Correlation 解決:
「這些可疑行為是不是同一件事情?」
兩者的目的並不一樣。
單一 Alert:
python3 -c
可能只是 False Positive。
但是:
python3 -c
↓
Interpreter Network Activity
↓
External Connection
可信度就會明顯提高。
因此:
Alert ≠ Incident
例如:
Execution
↓
Network
↓
Persistence
通常比:
三個互不相關的 Alert
更值得注意。
因此事件順序本身也是重要的 Detection Context。
一開始很容易覺得:
偵測越多
=
Alert 越多
=
系統越厲害
但實際上並不是。
如果一直出現:
Alert
Alert
Alert
Alert
Alert
最後分析人員可能直接忽略。
因此真正重要的是:
High Quality Alert
而不是:
High Quantity Alert
例如:
python3 -c
單獨出現可能只是正常行為。
但是:
python3 -c
+
External Connection
+
Interpreter Network Activity
三個行為在短時間內由同一 Process 產生,就更值得注意。
所以 Correlation 的價值之一,就是利用:
Context
增加判斷可信度。
Day14 是 BehaviorGuard 很重要的一個轉折點。
前面的系統主要是在做:
Process Detection
Command Detection
Network Detection
File Detection
Persistence Detection
也就是:
找到一個可疑行為,就產生一個 Alert。
今天則開始進入:
Correlation
把不同的 Alert 串成:
Attack Chain
目前已經能透過:
Time Window
+
PID
+
Correlation Rule
+
Rule Sequence
+
Duplicate Suppression
進行第一版事件關聯。
最後也成功建立:
BG-CMD-002
Python Inline Code Execution
↓
BG-NET-005
Shell or Interpreter Network Activity
↓
BG-NET-003
Unexpected External Connection
↓
BG-CORR-005
Suspicious Execution and Network Activity
↓
CRITICAL
BehaviorGuard 已經開始從:
「看到可疑行為」
逐漸進入:
「理解可疑行為之間的關係」
下一步 Day15,我準備繼續做:
加入:
PID
PPID
Parent Process
Ancestor Process
讓 BehaviorGuard 即使遇到:
bash
PID 7000
↓
python3
PID 7010
↓
nc
PID 7020
這種不同 PID 的情況,也能判斷:
它們可能其實屬於同一條攻擊鏈。
這也是下一階段要解決的問題。