iT邦幫忙

2026 iThome 鐵人賽

DAY 22
0
IT Operation

系統工程師的 30 天自動化維運實戰:PowerShell × AD × Windows Server系列 第 22 篇

Day 22|PowerShell 自動通知:巡檢出現 Warning / Critical 時主動發送 Email

  • 分享至 

  • xImage
  •  

Day 20 我們把 PowerShell Script 放進:
Task Scheduler

Day 21 又把:
CSV

整理成:
HTML Dashboard

目前整個流程已經可以做到:
06:00
↓
自動巡檢 Server / AD
↓
產生 CSV
↓
產生 HTML
↓
寫入 Log

但還有一個很現實的問題:
報表做得再漂亮,如果今天沒有人打開它,Critical 還是沒有人看到。

所以今天要再往前一步:
Notification
讓整套 Automation 變成:
Collect
↓
Analyze
↓
Report
↓
Detect Problem
↓
Notify Human

也就是:
平常不要一直吵我,真的有異常再主動叫我。

今天會先用 Email 做示範,並建立:
Healthy
→ 只產 Report

Warning
→ 發通知

Critical
→ 發高優先通知

Unknown / No Data
→ 也要通知

最後流程會變成:
Task Scheduler
↓
Health Check
↓
Status Analysis
↓
HTML Report
↓
Warning / Critical?
│
├── No
│ └── 結束
│
└── Yes
↓
Generate Alert
↓
Send Email

第一件事:不是所有狀況都要寄信
假設每天有 50 台 Server。
全部正常:
SERVER01 Healthy
SERVER02 Healthy
SERVER03 Healthy
...
SERVER50 Healthy

如果每天還寄:
[INFO] All Servers Healthy

久了很容易變成:
收件者看到 PowerShell Mail 就直接忽略。

這就是:
Alert Fatigue

所以我的第一版規則會是:
Healthy
→ 不主動通知

Warning
→ 通知

Critical
→ 通知

Unknown
→ 通知

No Data
→ 通知

尤其:
Unknown

不能忽略。
因為:
SERVER01 = Healthy

代表:
我有取得資料,而且目前正常。

但:
SERVER01 = Unknown

代表:
我根本不知道它現在正常不正常。

先找出異常 Server
假設前面的 Health Check 已經得到:
$Results

裡面有:
ComputerName
CPUStatus
MemoryStatus
DiskStatus
ServiceStatus
OverallStatus

可以:
$ProblemServers = @(
$Results |
Where-Object {

    $_.OverallStatus -in @(
        "Warning"
        "Critical"
        "Unknown"
    )

}

)

檢查:
$ProblemServers.Count

如果:
0

就代表:
目前沒有需要主動通知的 Server

不要只判斷 Critical
假設:
SERVER01
Connection = Failed
OverallStatus = Unknown

如果 Script 只寫:
Where-Object {
$_.OverallStatus -eq "Critical"
}

SERVER01 就完全不會被通知。
但:
監控不到本身就是值得處理的狀況。

所以今天會把:
Warning
Critical
Unknown

都列入通知候選。
先計算 Severity
例如:
$CriticalCount = @(
$Results |
Where-Object {
$_.OverallStatus -eq "Critical"
}
).Count

$WarningCount = @(
$Results |
Where-Object {
$_.OverallStatus -eq "Warning"
}
).Count

$UnknownCount = @(
$Results |
Where-Object {
$_.OverallStatus -eq "Unknown"
}
).Count

然後決定整份通知的最高 Severity:
if ($CriticalCount -gt 0) {

$AlertLevel = "CRITICAL"

}
elseif ($UnknownCount -gt 0) {

$AlertLevel = "UNKNOWN"

}
elseif ($WarningCount -gt 0) {

$AlertLevel = "WARNING"

}
else {

$AlertLevel = "HEALTHY"

}

現在我們就有:
整體 Alert Level

Email Subject 不要只寫「Server Report」
如果收到:
Subject:
Server Report

你還要打開才知道有沒有問題。
比較實用:
[CRITICAL] Infrastructure Health Check - 1 Critical / 2 Warning

例如 PowerShell:
$Subject =
"[$AlertLevel] Infrastructure Health Check - " +
"$CriticalCount Critical / " +
"$WarningCount Warning / " +
"$UnknownCount Unknown"

如果今天:
Critical = 1
Warning = 2
Unknown = 0

Subject 就是:
[CRITICAL] Infrastructure Health Check - 1 Critical / 2 Warning / 0 Unknown

收件者甚至不用打開 Email,就知道今天有事情。
Email 不要只寄一句「有錯誤」
例如:
Server has errors.

這種通知幫助不大。
至少應該包含:
哪台 Server?
什麼 Status?
CPU?
Memory?
Disk?
Service?
什麼時間?

所以我們先建立:
$AlertData =
$ProblemServers |
Select-Object `
ComputerName,
Connection,
CPUUsage,
CPUStatus,
MemoryUsage,
MemoryStatus,
DiskStatus,
ServiceStatus,
UptimeDays,
OverallStatus

用 ConvertTo-Html 做 Email Body
Day 21 已經學過:
ConvertTo-Html

今天直接重用。
$ProblemTable =
$AlertData |
ConvertTo-Html `
-Fragment

它會產生一段:

建立 Email Body
例如:
$GeneratedTime =
Get-Date `
-Format "yyyy-MM-dd HH:mm:ss"

然後:
$EmailBody = @"

$ProblemTable

"@

這樣寄出去的 Email 就不只是純文字。
Email 也可以加入簡單 CSS
例如:
$EmailCss = @"

"@

最後:
$EmailBody = @"

$EmailCss

$ProblemTable

"@

Send-MailMessage
Windows PowerShell 裡可以看到:
Send-MailMessage

例如最簡單:
Send-MailMessage -From "automation@contoso.com"
-To "it-ops@contoso.com" -Subject $Subject
-Body $EmailBody -BodyAsHtml
-SmtpServer "smtp-relay.contoso.com"

如果公司有內部 SMTP Relay,這是很常見、也很容易理解的 Lab 範例。
但 Send-MailMessage 不適合當長期新架構的唯一選擇
這裡特別提醒。
Send-MailMessage 屬於比較舊的 PowerShell 郵件方式,Microsoft 已將它標示為 obsolete。
所以今天使用它主要是因為:
容易理解
Windows PowerShell 5.1 常見
很適合學 SMTP Notification 流程

新的正式環境則應優先依公司的郵件架構選擇,例如:
Internal SMTP Relay

Microsoft Graph
(Microsoft 365 環境)

企業既有 Notification Service

Teams / Workflow / Webhook

也就是今天真正要學的是:
Notification Workflow

而不是死背某一個寄信 Cmdlet。
不要把 SMTP Password 寫在 Script
不要:
$Password = "MyMailPassword123"

也不要:
SMTP Username / Password

放在:
.ps1
.csv
.txt

如果公司的 SMTP Relay 是依:
來源 IP
服務帳號
內部 Relay Policy

控制,就按照公司的既有方案。
如果需要 Authentication,應使用公司核准的 Credential / Secret 管理方式。
Email Recipient 也可以放 Configuration
不要每一支 Script 都:
-To "ted@contoso.com"

寫死。
可以建立:
NotificationConfig.csv

例如:
AlertType,Recipient
Infrastructure,it-ops@contoso.com
ActiveDirectory,ad-admins@contoso.com
Security,security@contoso.com

未來:
$NotificationConfig =
Import-Csv `
"$PSScriptRoot\Config\NotificationConfig.csv"

這就延續 Day 12 的:
Configuration 與 Logic 分離。

先做 ShouldSendAlert
我不想主程式裡到處寫:
if (...)

可以:
$ShouldSendAlert =
$ProblemServers.Count -gt 0

然後:
if ($ShouldSendAlert) {

# Send Alert

}
else {

Write-Log `
    "No alert notification required."

}

流程就很清楚。
HTML Report 可以直接當 Attachment
Day 21 已經產生:
Daily_Health_Report_20260930.html

所以 Email 可以:
Send-MailMessage -From "automation@contoso.com"
-To "it-ops@contoso.com" -Subject $Subject
-Body $EmailBody -BodyAsHtml
-SmtpServer "smtp-relay.contoso.com" `
-Attachments $HtmlFile

收件者就可以:
先看 Email Summary
↓
需要更多資訊
↓
開完整 HTML Report

但 Attachment 不是唯一方式
如果 Report 已經放在內部:
Web Server
SharePoint
File Server
Internal Portal

更好的方式可能是寄:
Report Link

而不是每天附:
HTML
CSV
CSV
CSV

Email 越來越大。
所以架構可以是:
Email
│
├── Summary
├── Critical Items
└── Report URL

如果 Report 放在 File Share
例如:
\FILESERVER01\ITReports\

要注意:
收件者看到 UNC Path,不代表一定有權限。

而且從手機、外部環境可能完全打不開。
所以如果真的要給較多人看,內部 HTTP / Portal 往往會比檔案路徑好用。
今天先用 Attachment 做 Lab。
Notification 失敗要不要算整個 Health Check Failed?
這是一個很有意思的問題。
假設:
Server Health Check
→ Success

發現 Critical
→ Success

HTML Report
→ Success

Email
→ Failed

這時候:
Server Health Check 本身成功了。

但是:
Human Notification 沒有完成。

所以我不會把它當:
Everything Success

比較適合:
Automation Overall

Partial

因為主要資料有收集成功,但 Notification Channel 出問題。
加入 NotificationStatus
例如:
$NotificationStatus =
"NotRequired"

如果有異常:
$NotificationStatus =
"Pending"

寄信成功:
$NotificationStatus =
"Success"

寄信失敗:
$NotificationStatus =
"Failed"

這跟 Day 19 的:
DisableStatus
MoveOUStatus
VerificationStatus

是一樣的設計。
寄信也要 try/catch
例如:
try {

Send-MailMessage `
    -From "automation@contoso.com" `
    -To "it-ops@contoso.com" `
    -Subject $Subject `
    -Body $EmailBody `
    -BodyAsHtml `
    -SmtpServer "smtp-relay.contoso.com" `
    -ErrorAction Stop


$NotificationStatus =
    "Success"


Write-Log `
    "Alert email sent successfully."

}
catch {

$NotificationStatus =
    "Failed"


Write-Log `
    -Message "Email notification failed: $($_.Exception.Message)" `
    -Level "ERROR"

}

不要:
Send-MailMessage
↓
沒報錯就不記錄

因為排程之後需要知道:
今天到底有沒有通知成功?

如果 Email 失敗,還有 Log
這就是為什麼:
Notification

不能取代:
Log

整套架構應該是:
Log
→ 系統自己的執行紀錄

CSV
→ Raw Data

HTML
→ Human-readable Report

Email
→ 主動通知

四個功能不同。
不要只通知 Critical
例如:
Disk = 82%
Warning

今天沒有處理。
一週後可能:
Disk = 97%
Critical

所以 Warning 的目的就是:
在真正出事以前先看到趨勢。

但 Warning 太多也會造成 Alert Fatigue。
這時候就需要後面慢慢調整:
Threshold
Suppress
Ignore Rule
Maintenance Window
Known Issue

今天先不做得太複雜。
建立問題摘要
我們甚至可以在 Email 裡直接產:
SERVER03
Overall: Critical
CPU: 94%
Disk: Critical

SERVER05
Overall: Warning
Memory: 83%

例如:
$ProblemSummary =
foreach ($Server in $ProblemServers) {

    [PSCustomObject]@{

        Server =
            $Server.ComputerName

        Overall =
            $Server.OverallStatus

        CPU =
            "$($Server.CPUUsage)%"

        CPUStatus =
            $Server.CPUStatus

        Memory =
            "$($Server.MemoryUsage)%"

        MemoryStatus =
            $Server.MemoryStatus

        Disk =
            $Server.DiskStatus

        Service =
            $Server.ServiceStatus
    }
}

再:
$ProblemTable =
$ProblemSummary |
ConvertTo-Html `
-Fragment

這樣 Email 就很容易看。
Critical 排最前面
跟 Day 21 一樣:
$ProblemServers =
$ProblemServers |
Sort-Object @{
Expression = {

        switch (
            $_.OverallStatus
        ) {

            "Critical" {
                1
            }

            "Unknown" {
                2
            }

            "Warning" {
                3
            }

            default {
                4
            }
        }
    }
}

讓最重要的問題先出現。
Active Directory 也可以放通知
不只 Server。
例如 Day 14:
Inactive Users = 25

通常不一定需要立刻寄:
Critical

但如果 Day 16 發現:
Disabled User
還留在 Server-Admins

這可能更值得主動通知。
所以未來 Notification 可以分:
Infrastructure Alert

AD Security / Access Review

Daily Audit Summary

不要所有東西都塞在同一封 Email。
不同通知可以有不同 Severity
例如:
Server CPU Warning
→ Warning

Disk Critical
→ Critical

Connection Unknown
→ Unknown

Disabled User in Sensitive Group
→ Review

Inactive User
→ Audit

這樣未來 Report / Email 會更接近:
事件分類。

而不是全部都叫:
Error

今天完整 Notification 範例
以下接 Day 21 的 $Results 與 $HtmlFile。

==========================================

Infrastructure Alert Notification

Day 22

==========================================

==========================================

Configuration

==========================================

$MailFrom =
"automation@contoso.com"

$MailTo =
"it-ops@contoso.com"

$SmtpServer =
"smtp-relay.contoso.com"

==========================================

Detect Problems

==========================================

$ProblemServers = @(
$Results |
Where-Object {

    $_.OverallStatus -in @(
        "Warning",
        "Critical",
        "Unknown"
    )

}

)

$CriticalCount = @(
$ProblemServers |
Where-Object {
$_.OverallStatus -eq "Critical"
}
).Count

$WarningCount = @(
$ProblemServers |
Where-Object {
$_.OverallStatus -eq "Warning"
}
).Count

$UnknownCount = @(
$ProblemServers |
Where-Object {
$_.OverallStatus -eq "Unknown"
}
).Count

==========================================

Alert Level

==========================================

if ($CriticalCount -gt 0) {

$AlertLevel =
    "CRITICAL"

}
elseif ($UnknownCount -gt 0) {

$AlertLevel =
    "UNKNOWN"

}
elseif ($WarningCount -gt 0) {

$AlertLevel =
    "WARNING"

}
else {

$AlertLevel =
    "HEALTHY"

}

==========================================

Notification State

==========================================

$NotificationStatus =
"NotRequired"

==========================================

Only Notify on Problem

==========================================

if ($ProblemServers.Count -gt 0) {

$NotificationStatus =
    "Pending"


# ======================================
# Sort Problems
# ======================================

$ProblemServers =
    $ProblemServers |
    Sort-Object @{
        Expression = {

            switch (
                $_.OverallStatus
            ) {

                "Critical" {
                    1
                }

                "Unknown" {
                    2
                }

                "Warning" {
                    3
                }

                default {
                    4
                }
            }
        }
    }


# ======================================
# Email Table
# ======================================

$ProblemSummary =
    $ProblemServers |
    Select-Object `
        ComputerName,
        Connection,
        CPUUsage,
        CPUStatus,
        MemoryUsage,
        MemoryStatus,
        DiskStatus,
        ServiceStatus,
        OverallStatus


$ProblemTable =
    $ProblemSummary |
    ConvertTo-Html `
        -Fragment


# ======================================
# Subject
# ======================================

$Subject =
    "[$AlertLevel] Infrastructure Health - " +
    "$CriticalCount Critical / " +
    "$WarningCount Warning / " +
    "$UnknownCount Unknown"


# ======================================
# CSS
# ======================================

$EmailCss = @"

"@

# ======================================
# Body
# ======================================

$GeneratedTime =
    Get-Date `
        -Format "yyyy-MM-dd HH:mm:ss"


$EmailBody = @"

$EmailCss

$ProblemTable

Please review the attached
infrastructure health report
for additional details.

"@

# ======================================
# Send
# ======================================

try {

    $MailParameters = @{

        From =
            $MailFrom

        To =
            $MailTo

        Subject =
            $Subject

        Body =
            $EmailBody

        BodyAsHtml =
            $true

        SmtpServer =
            $SmtpServer

        ErrorAction =
            "Stop"
    }


    if (
        Test-Path $HtmlFile
    ) {

        $MailParameters.Attachments =
            $HtmlFile
    }


    Send-MailMessage `
        @MailParameters


    $NotificationStatus =
        "Success"


    Write-Log `
        "Alert email sent successfully."

}
catch {

    $NotificationStatus =
        "Failed"


    Write-Log `
        -Message "Alert email failed: $($_.Exception.Message)" `
        -Level "ERROR"
}

}
else {

Write-Log `
    "All servers healthy. Notification not required."

}

==========================================

Display

==========================================

Write-Host ""

Write-Host `
"Alert Level: $AlertLevel"

Write-Host `
"Notification: $NotificationStatus"

PowerShell Splatting
上面還出現一個很好用的技巧:
$MailParameters = @{
From = ...
To = ...
Subject = ...
}

然後:
Send-MailMessage @MailParameters

這叫:
Splatting
如果原本寫:
Send-MailMessage -From ...
-To ... -Subject ...
-Body ... -BodyAsHtml
-SmtpServer ... `
-Attachments ...

參數很多時會越來越長。
Splatting 可以整理成:
Parameter
↓
Hashtable
↓
Command

後面的 PowerShell Automation 會很常用到。
Attachment 存在才加入
這裡我們沒有直接:
-Attachments $HtmlFile

而是:
if (Test-Path $HtmlFile) {

$MailParameters.Attachments =
    $HtmlFile

}

因為:
HTML Report 不存在,不應該讓整封 Alert 因 Attachment 找不到而直接失敗。

這也是:
Graceful Degradation

的概念。
理想狀態:
Email + HTML

如果 HTML 失敗:
至少 Email Summary 還能出去

Notification 也要進入 Overall Status
假設:
HealthCheck = Success
Report = Success
Email = Failed

最後可以:
if (
$NotificationStatus -eq "Failed"
) {

$HasPartialError =
    $true

}

然後延續 Day 20:
if ($HasFatalError) {

exit 1

}
elseif ($HasPartialError) {

exit 2

}
else {

exit 0

}

現在:
Exit 0
→ 完整成功

Exit 1
→ Fatal Failure

Exit 2
→ 有工作完成,但存在部分問題

Task Scheduler + Email
現在整體流程就變成:
06:00
│
▼
Task Scheduler
│
▼
PowerShell
│
├── Server Health
├── AD Audit
├── CSV
├── HTML
│
▼
Status Analysis
│
├── Healthy
│ ↓
│ No Mail
│
├── Warning
│ ↓
│ Send Alert
│
├── Critical
│ ↓
│ Send Alert
│
└── Unknown
↓
Send Alert

到這裡,工程師終於不需要:
每天自己想起來去看 Report。

但 Email Alert 有一個大問題:重複通知
假設:
SERVER03 Disk = 95%

星期一:
Critical
→ Mail

星期二:
還是 Critical
→ Mail

星期三:
還是 Critical
→ Mail

如果問題一直沒有修:
每天同樣一封

最後還是會變成 Alert Fatigue。
未來可以加入 State
例如保存:
昨天的狀態

今天比較:
昨天 Healthy
今天 Critical
→ 新問題
→ Notify

昨天 Critical
今天 Critical
→ Existing Issue

昨天 Critical
今天 Healthy
→ Recovery
→ Notify

這就開始從:
Stateless Script

走向:
Stateful Monitoring

後面可以再做。
Recovery Notification 其實也很重要
如果昨天收到:
SERVER03 Disk Critical

今天工程師處理完:
SERVER03 Healthy

最好有:
[RECOVERY] SERVER03 returned to Healthy

否則收件者只知道:
昨天壞了。

不知道:
現在好了沒有。

這就是監控系統常見的:
Alert
↓
Recovery

概念。
Email 不應該成為唯一 Alert Channel
正式環境還可能需要:
Email
Teams
Slack
Pager / On-call
Ticket System
Monitoring Platform

不同 Severity 可以走不同 Channel。
例如:
Warning
→ Email

Critical
→ Email + Teams

P1
→ Monitoring / On-call

今天先把 Notification Logic 建起來。
Channel 後面可以替換。
Day 22 小結
今天我們終於把:
PowerShell

從:
會執行
↓
會產資料
↓
會做報表

再往前推進成:
會判斷結果
↓
發現異常
↓
主動通知人

主要流程是:
Health Check
↓
Problem Detection
↓
Severity
↓
Warning / Critical / Unknown
↓
HTML Summary
↓
Notification
↓
Notification Status

今天也開始把:
Notification

當成 Automation Workflow 裡的一個正式 Step。
也就是:
HealthCheckStatus

ReportStatus

NotificationStatus

OverallStatus

而不是:
寄信只是最後順手加一行。

另外今天幾個重要觀念是:
Healthy 可以不通知,但 Unknown 不能等同 Healthy。

因為 No Data / Connection Failure 本身就是值得 Review 的資訊。
第二個:
Alert 應該包含足夠資訊,讓人不用先登入 Server 才知道發生什麼事。

至少應該有:
Server
Severity
CPU
Memory
Disk
Service
Time

第三個:
Send-MailMessage 可以拿來理解 SMTP Automation,但新正式環境不應只依賴這個舊式 Cmdlet。

應該依公司的:
SMTP Relay
Microsoft 365 / Graph
Teams
Notification Platform

選擇真正的通知方式。
最後則是:
Notification 自己也可能失敗。

所以:
Health Check Success
+
Email Failed

不能直接算:
Everything Success

而應該留下:
Partial

供後續處理。
現在我們的 Automation 已經變成:
Schedule
↓
Collect
↓
Analyze
↓
Report
↓
Notify
↓
Human Action

開始有一個完整「維運閉環」的雛形。
Day 23 預告
Day 23|PowerShell 狀態追蹤:不要每天重複告警,只通知「新異常」與「恢復」
Day 22 還留下一個很重要的問題。
假設:
SERVER03
Disk = Critical

今天寄:
Critical Alert

明天還沒修:
又寄一次

後天:
又寄一次

久了就變成:
告警很多,但沒有人再看。

所以 Day 23 我們會第一次加入:
Previous State

例如保存:
SERVER01 = Healthy
SERVER02 = Warning
SERVER03 = Critical

下一次巡檢後比較:
Previous
vs
Current

判斷:
Healthy → Critical
= New Alert

Warning → Critical
= Escalated

Critical → Critical
= Existing Issue

Critical → Healthy
= Recovery

Healthy → Healthy
= No Notification

最後讓通知邏輯從:
現在有問題嗎?

升級成:
「這次跟上次相比,到底發生了什麼改變?」

Day 23 會正式讓我們的 Script 開始有 State,也會更接近真正 Monitoring System 的事件處理方式。


上一篇
Day 21|PowerShell HTML 報表:把冷冰冰的 CSV 變成一眼看懂的巡檢 Dashboard
系列文
系統工程師的 30 天自動化維運實戰:PowerShell × AD × Windows Server 共 22 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言