Day 20 我們把 PowerShell Script 放進:
Task Scheduler
Day 21 又把:
CSV
整理成:
HTML Dashboard
目前整個流程已經可以做到:
06:00
↓
自動巡檢 Server / AD
↓
產生 CSV
↓
產生 HTML
↓
寫入 Log
但還有一個很現實的問題:
報表做得再漂亮,如果今天沒有人打開它,Critical 還是沒有人看到。
所以今天要再往前一步:
Notification
讓整套 Automation 變成:
Collect
↓
Analyze
↓
Report
↓
Detect Problem
↓
Notify Human
也就是:
平常不要一直吵我,真的有異常再主動叫我。
今天會先用 Email 做示範,並建立:
Healthy
→ 只產 Report
Warning
→ 發通知
Critical
→ 發高優先通知
Unknown / No Data
→ 也要通知
最後流程會變成:
Task Scheduler
↓
Health Check
↓
Status Analysis
↓
HTML Report
↓
Warning / Critical?
│
├── No
│ └── 結束
│
└── Yes
↓
Generate Alert
↓
Send Email
第一件事:不是所有狀況都要寄信
假設每天有 50 台 Server。
全部正常:
SERVER01 Healthy
SERVER02 Healthy
SERVER03 Healthy
...
SERVER50 Healthy
如果每天還寄:
[INFO] All Servers Healthy
久了很容易變成:
收件者看到 PowerShell Mail 就直接忽略。
這就是:
Alert Fatigue
所以我的第一版規則會是:
Healthy
→ 不主動通知
Warning
→ 通知
Critical
→ 通知
Unknown
→ 通知
No Data
→ 通知
尤其:
Unknown
不能忽略。
因為:
SERVER01 = Healthy
代表:
我有取得資料,而且目前正常。
但:
SERVER01 = Unknown
代表:
我根本不知道它現在正常不正常。
先找出異常 Server
假設前面的 Health Check 已經得到:
$Results
裡面有:
ComputerName
CPUStatus
MemoryStatus
DiskStatus
ServiceStatus
OverallStatus
可以:
$ProblemServers = @(
$Results |
Where-Object {
$_.OverallStatus -in @(
"Warning"
"Critical"
"Unknown"
)
}
)
檢查:
$ProblemServers.Count
如果:
0
就代表:
目前沒有需要主動通知的 Server
不要只判斷 Critical
假設:
SERVER01
Connection = Failed
OverallStatus = Unknown
如果 Script 只寫:
Where-Object {
$_.OverallStatus -eq "Critical"
}
SERVER01 就完全不會被通知。
但:
監控不到本身就是值得處理的狀況。
所以今天會把:
Warning
Critical
Unknown
都列入通知候選。
先計算 Severity
例如:
$CriticalCount = @(
$Results |
Where-Object {
$_.OverallStatus -eq "Critical"
}
).Count
$WarningCount = @(
$Results |
Where-Object {
$_.OverallStatus -eq "Warning"
}
).Count
$UnknownCount = @(
$Results |
Where-Object {
$_.OverallStatus -eq "Unknown"
}
).Count
然後決定整份通知的最高 Severity:
if ($CriticalCount -gt 0) {
$AlertLevel = "CRITICAL"
}
elseif ($UnknownCount -gt 0) {
$AlertLevel = "UNKNOWN"
}
elseif ($WarningCount -gt 0) {
$AlertLevel = "WARNING"
}
else {
$AlertLevel = "HEALTHY"
}
現在我們就有:
整體 Alert Level
Email Subject 不要只寫「Server Report」
如果收到:
Subject:
Server Report
你還要打開才知道有沒有問題。
比較實用:
[CRITICAL] Infrastructure Health Check - 1 Critical / 2 Warning
例如 PowerShell:
$Subject =
"[$AlertLevel] Infrastructure Health Check - " +
"$CriticalCount Critical / " +
"$WarningCount Warning / " +
"$UnknownCount Unknown"
如果今天:
Critical = 1
Warning = 2
Unknown = 0
Subject 就是:
[CRITICAL] Infrastructure Health Check - 1 Critical / 2 Warning / 0 Unknown
收件者甚至不用打開 Email,就知道今天有事情。
Email 不要只寄一句「有錯誤」
例如:
Server has errors.
這種通知幫助不大。
至少應該包含:
哪台 Server?
什麼 Status?
CPU?
Memory?
Disk?
Service?
什麼時間?
所以我們先建立:
$AlertData =
$ProblemServers |
Select-Object `
ComputerName,
Connection,
CPUUsage,
CPUStatus,
MemoryUsage,
MemoryStatus,
DiskStatus,
ServiceStatus,
UptimeDays,
OverallStatus
用 ConvertTo-Html 做 Email Body
Day 21 已經學過:
ConvertTo-Html
今天直接重用。
$ProblemTable =
$AlertData |
ConvertTo-Html `
-Fragment
它會產生一段:
建立 Email Body
例如:
$GeneratedTime =
Get-Date `
-Format "yyyy-MM-dd HH:mm:ss"
然後:
$EmailBody = @"
$ProblemTable
"@
這樣寄出去的 Email 就不只是純文字。
Email 也可以加入簡單 CSS
例如:
$EmailCss = @"
"@
最後:
$EmailBody = @"
$EmailCss
$ProblemTable
"@
Send-MailMessage
Windows PowerShell 裡可以看到:
Send-MailMessage
例如最簡單:
Send-MailMessage -From "automation@contoso.com"
-To "it-ops@contoso.com" -Subject $Subject
-Body $EmailBody -BodyAsHtml
-SmtpServer "smtp-relay.contoso.com"
如果公司有內部 SMTP Relay,這是很常見、也很容易理解的 Lab 範例。
但 Send-MailMessage 不適合當長期新架構的唯一選擇
這裡特別提醒。
Send-MailMessage 屬於比較舊的 PowerShell 郵件方式,Microsoft 已將它標示為 obsolete。
所以今天使用它主要是因為:
容易理解
Windows PowerShell 5.1 常見
很適合學 SMTP Notification 流程
新的正式環境則應優先依公司的郵件架構選擇,例如:
Internal SMTP Relay
Microsoft Graph
(Microsoft 365 環境)
企業既有 Notification Service
Teams / Workflow / Webhook
也就是今天真正要學的是:
Notification Workflow
而不是死背某一個寄信 Cmdlet。
不要把 SMTP Password 寫在 Script
不要:
$Password = "MyMailPassword123"
也不要:
SMTP Username / Password
放在:
.ps1
.csv
.txt
如果公司的 SMTP Relay 是依:
來源 IP
服務帳號
內部 Relay Policy
控制,就按照公司的既有方案。
如果需要 Authentication,應使用公司核准的 Credential / Secret 管理方式。
Email Recipient 也可以放 Configuration
不要每一支 Script 都:
-To "ted@contoso.com"
寫死。
可以建立:
NotificationConfig.csv
例如:
AlertType,Recipient
Infrastructure,it-ops@contoso.com
ActiveDirectory,ad-admins@contoso.com
Security,security@contoso.com
未來:
$NotificationConfig =
Import-Csv `
"$PSScriptRoot\Config\NotificationConfig.csv"
這就延續 Day 12 的:
Configuration 與 Logic 分離。
先做 ShouldSendAlert
我不想主程式裡到處寫:
if (...)
可以:
$ShouldSendAlert =
$ProblemServers.Count -gt 0
然後:
if ($ShouldSendAlert) {
# Send Alert
}
else {
Write-Log `
"No alert notification required."
}
流程就很清楚。
HTML Report 可以直接當 Attachment
Day 21 已經產生:
Daily_Health_Report_20260930.html
所以 Email 可以:
Send-MailMessage -From "automation@contoso.com"
-To "it-ops@contoso.com" -Subject $Subject
-Body $EmailBody -BodyAsHtml
-SmtpServer "smtp-relay.contoso.com" `
-Attachments $HtmlFile
收件者就可以:
先看 Email Summary
↓
需要更多資訊
↓
開完整 HTML Report
但 Attachment 不是唯一方式
如果 Report 已經放在內部:
Web Server
SharePoint
File Server
Internal Portal
更好的方式可能是寄:
Report Link
而不是每天附:
HTML
CSV
CSV
CSV
Email 越來越大。
所以架構可以是:
Email
│
├── Summary
├── Critical Items
└── Report URL
如果 Report 放在 File Share
例如:
\FILESERVER01\ITReports\
要注意:
收件者看到 UNC Path,不代表一定有權限。
而且從手機、外部環境可能完全打不開。
所以如果真的要給較多人看,內部 HTTP / Portal 往往會比檔案路徑好用。
今天先用 Attachment 做 Lab。
Notification 失敗要不要算整個 Health Check Failed?
這是一個很有意思的問題。
假設:
Server Health Check
→ Success
發現 Critical
→ Success
HTML Report
→ Success
Email
→ Failed
這時候:
Server Health Check 本身成功了。
但是:
Human Notification 沒有完成。
所以我不會把它當:
Everything Success
Partial
因為主要資料有收集成功,但 Notification Channel 出問題。
加入 NotificationStatus
例如:
$NotificationStatus =
"NotRequired"
如果有異常:
$NotificationStatus =
"Pending"
寄信成功:
$NotificationStatus =
"Success"
寄信失敗:
$NotificationStatus =
"Failed"
這跟 Day 19 的:
DisableStatus
MoveOUStatus
VerificationStatus
是一樣的設計。
寄信也要 try/catch
例如:
try {
Send-MailMessage `
-From "automation@contoso.com" `
-To "it-ops@contoso.com" `
-Subject $Subject `
-Body $EmailBody `
-BodyAsHtml `
-SmtpServer "smtp-relay.contoso.com" `
-ErrorAction Stop
$NotificationStatus =
"Success"
Write-Log `
"Alert email sent successfully."
}
catch {
$NotificationStatus =
"Failed"
Write-Log `
-Message "Email notification failed: $($_.Exception.Message)" `
-Level "ERROR"
}
不要:
Send-MailMessage
↓
沒報錯就不記錄
因為排程之後需要知道:
今天到底有沒有通知成功?
如果 Email 失敗,還有 Log
這就是為什麼:
Notification
不能取代:
Log
整套架構應該是:
Log
→ 系統自己的執行紀錄
CSV
→ Raw Data
HTML
→ Human-readable Report
Email
→ 主動通知
四個功能不同。
不要只通知 Critical
例如:
Disk = 82%
Warning
今天沒有處理。
一週後可能:
Disk = 97%
Critical
所以 Warning 的目的就是:
在真正出事以前先看到趨勢。
但 Warning 太多也會造成 Alert Fatigue。
這時候就需要後面慢慢調整:
Threshold
Suppress
Ignore Rule
Maintenance Window
Known Issue
今天先不做得太複雜。
建立問題摘要
我們甚至可以在 Email 裡直接產:
SERVER03
Overall: Critical
CPU: 94%
Disk: Critical
SERVER05
Overall: Warning
Memory: 83%
例如:
$ProblemSummary =
foreach ($Server in $ProblemServers) {
[PSCustomObject]@{
Server =
$Server.ComputerName
Overall =
$Server.OverallStatus
CPU =
"$($Server.CPUUsage)%"
CPUStatus =
$Server.CPUStatus
Memory =
"$($Server.MemoryUsage)%"
MemoryStatus =
$Server.MemoryStatus
Disk =
$Server.DiskStatus
Service =
$Server.ServiceStatus
}
}
再:
$ProblemTable =
$ProblemSummary |
ConvertTo-Html `
-Fragment
這樣 Email 就很容易看。
Critical 排最前面
跟 Day 21 一樣:
$ProblemServers =
$ProblemServers |
Sort-Object @{
Expression = {
switch (
$_.OverallStatus
) {
"Critical" {
1
}
"Unknown" {
2
}
"Warning" {
3
}
default {
4
}
}
}
}
讓最重要的問題先出現。
Active Directory 也可以放通知
不只 Server。
例如 Day 14:
Inactive Users = 25
通常不一定需要立刻寄:
Critical
但如果 Day 16 發現:
Disabled User
還留在 Server-Admins
這可能更值得主動通知。
所以未來 Notification 可以分:
Infrastructure Alert
AD Security / Access Review
Daily Audit Summary
不要所有東西都塞在同一封 Email。
不同通知可以有不同 Severity
例如:
Server CPU Warning
→ Warning
Disk Critical
→ Critical
Connection Unknown
→ Unknown
Disabled User in Sensitive Group
→ Review
Inactive User
→ Audit
這樣未來 Report / Email 會更接近:
事件分類。
而不是全部都叫:
Error
今天完整 Notification 範例
以下接 Day 21 的 $Results 與 $HtmlFile。
$MailFrom =
"automation@contoso.com"
$MailTo =
"it-ops@contoso.com"
$SmtpServer =
"smtp-relay.contoso.com"
$ProblemServers = @(
$Results |
Where-Object {
$_.OverallStatus -in @(
"Warning",
"Critical",
"Unknown"
)
}
)
$CriticalCount = @(
$ProblemServers |
Where-Object {
$_.OverallStatus -eq "Critical"
}
).Count
$WarningCount = @(
$ProblemServers |
Where-Object {
$_.OverallStatus -eq "Warning"
}
).Count
$UnknownCount = @(
$ProblemServers |
Where-Object {
$_.OverallStatus -eq "Unknown"
}
).Count
if ($CriticalCount -gt 0) {
$AlertLevel =
"CRITICAL"
}
elseif ($UnknownCount -gt 0) {
$AlertLevel =
"UNKNOWN"
}
elseif ($WarningCount -gt 0) {
$AlertLevel =
"WARNING"
}
else {
$AlertLevel =
"HEALTHY"
}
$NotificationStatus =
"NotRequired"
if ($ProblemServers.Count -gt 0) {
$NotificationStatus =
"Pending"
# ======================================
# Sort Problems
# ======================================
$ProblemServers =
$ProblemServers |
Sort-Object @{
Expression = {
switch (
$_.OverallStatus
) {
"Critical" {
1
}
"Unknown" {
2
}
"Warning" {
3
}
default {
4
}
}
}
}
# ======================================
# Email Table
# ======================================
$ProblemSummary =
$ProblemServers |
Select-Object `
ComputerName,
Connection,
CPUUsage,
CPUStatus,
MemoryUsage,
MemoryStatus,
DiskStatus,
ServiceStatus,
OverallStatus
$ProblemTable =
$ProblemSummary |
ConvertTo-Html `
-Fragment
# ======================================
# Subject
# ======================================
$Subject =
"[$AlertLevel] Infrastructure Health - " +
"$CriticalCount Critical / " +
"$WarningCount Warning / " +
"$UnknownCount Unknown"
# ======================================
# CSS
# ======================================
$EmailCss = @"
"@
# ======================================
# Body
# ======================================
$GeneratedTime =
Get-Date `
-Format "yyyy-MM-dd HH:mm:ss"
$EmailBody = @"
$EmailCss
$ProblemTable
Please review the attached
infrastructure health report
for additional details.
"@
# ======================================
# Send
# ======================================
try {
$MailParameters = @{
From =
$MailFrom
To =
$MailTo
Subject =
$Subject
Body =
$EmailBody
BodyAsHtml =
$true
SmtpServer =
$SmtpServer
ErrorAction =
"Stop"
}
if (
Test-Path $HtmlFile
) {
$MailParameters.Attachments =
$HtmlFile
}
Send-MailMessage `
@MailParameters
$NotificationStatus =
"Success"
Write-Log `
"Alert email sent successfully."
}
catch {
$NotificationStatus =
"Failed"
Write-Log `
-Message "Alert email failed: $($_.Exception.Message)" `
-Level "ERROR"
}
}
else {
Write-Log `
"All servers healthy. Notification not required."
}
Write-Host ""
Write-Host `
"Alert Level: $AlertLevel"
Write-Host `
"Notification: $NotificationStatus"
PowerShell Splatting
上面還出現一個很好用的技巧:
$MailParameters = @{
From = ...
To = ...
Subject = ...
}
然後:
Send-MailMessage @MailParameters
這叫:
Splatting
如果原本寫:
Send-MailMessage -From ...
-To ... -Subject ...
-Body ... -BodyAsHtml
-SmtpServer ... `
-Attachments ...
參數很多時會越來越長。
Splatting 可以整理成:
Parameter
↓
Hashtable
↓
Command
後面的 PowerShell Automation 會很常用到。
Attachment 存在才加入
這裡我們沒有直接:
-Attachments $HtmlFile
而是:
if (Test-Path $HtmlFile) {
$MailParameters.Attachments =
$HtmlFile
}
因為:
HTML Report 不存在,不應該讓整封 Alert 因 Attachment 找不到而直接失敗。
這也是:
Graceful Degradation
的概念。
理想狀態:
Email + HTML
如果 HTML 失敗:
至少 Email Summary 還能出去
Notification 也要進入 Overall Status
假設:
HealthCheck = Success
Report = Success
Email = Failed
最後可以:
if (
$NotificationStatus -eq "Failed"
) {
$HasPartialError =
$true
}
然後延續 Day 20:
if ($HasFatalError) {
exit 1
}
elseif ($HasPartialError) {
exit 2
}
else {
exit 0
}
現在:
Exit 0
→ 完整成功
Exit 1
→ Fatal Failure
Exit 2
→ 有工作完成,但存在部分問題
Task Scheduler + Email
現在整體流程就變成:
06:00
│
▼
Task Scheduler
│
▼
PowerShell
│
├── Server Health
├── AD Audit
├── CSV
├── HTML
│
▼
Status Analysis
│
├── Healthy
│ ↓
│ No Mail
│
├── Warning
│ ↓
│ Send Alert
│
├── Critical
│ ↓
│ Send Alert
│
└── Unknown
↓
Send Alert
到這裡,工程師終於不需要:
每天自己想起來去看 Report。
但 Email Alert 有一個大問題:重複通知
假設:
SERVER03 Disk = 95%
星期一:
Critical
→ Mail
星期二:
還是 Critical
→ Mail
星期三:
還是 Critical
→ Mail
如果問題一直沒有修:
每天同樣一封
最後還是會變成 Alert Fatigue。
未來可以加入 State
例如保存:
昨天的狀態
今天比較:
昨天 Healthy
今天 Critical
→ 新問題
→ Notify
昨天 Critical
今天 Critical
→ Existing Issue
昨天 Critical
今天 Healthy
→ Recovery
→ Notify
這就開始從:
Stateless Script
走向:
Stateful Monitoring
後面可以再做。
Recovery Notification 其實也很重要
如果昨天收到:
SERVER03 Disk Critical
今天工程師處理完:
SERVER03 Healthy
最好有:
[RECOVERY] SERVER03 returned to Healthy
否則收件者只知道:
昨天壞了。
不知道:
現在好了沒有。
這就是監控系統常見的:
Alert
↓
Recovery
概念。
Email 不應該成為唯一 Alert Channel
正式環境還可能需要:
Email
Teams
Slack
Pager / On-call
Ticket System
Monitoring Platform
不同 Severity 可以走不同 Channel。
例如:
Warning
→ Email
Critical
→ Email + Teams
P1
→ Monitoring / On-call
今天先把 Notification Logic 建起來。
Channel 後面可以替換。
Day 22 小結
今天我們終於把:
PowerShell
從:
會執行
↓
會產資料
↓
會做報表
再往前推進成:
會判斷結果
↓
發現異常
↓
主動通知人
主要流程是:
Health Check
↓
Problem Detection
↓
Severity
↓
Warning / Critical / Unknown
↓
HTML Summary
↓
Notification
↓
Notification Status
今天也開始把:
Notification
當成 Automation Workflow 裡的一個正式 Step。
也就是:
HealthCheckStatus
ReportStatus
NotificationStatus
OverallStatus
而不是:
寄信只是最後順手加一行。
另外今天幾個重要觀念是:
Healthy 可以不通知,但 Unknown 不能等同 Healthy。
因為 No Data / Connection Failure 本身就是值得 Review 的資訊。
第二個:
Alert 應該包含足夠資訊,讓人不用先登入 Server 才知道發生什麼事。
至少應該有:
Server
Severity
CPU
Memory
Disk
Service
Time
第三個:
Send-MailMessage 可以拿來理解 SMTP Automation,但新正式環境不應只依賴這個舊式 Cmdlet。
應該依公司的:
SMTP Relay
Microsoft 365 / Graph
Teams
Notification Platform
選擇真正的通知方式。
最後則是:
Notification 自己也可能失敗。
所以:
Health Check Success
+
Email Failed
不能直接算:
Everything Success
而應該留下:
Partial
供後續處理。
現在我們的 Automation 已經變成:
Schedule
↓
Collect
↓
Analyze
↓
Report
↓
Notify
↓
Human Action
開始有一個完整「維運閉環」的雛形。
Day 23 預告
Day 23|PowerShell 狀態追蹤:不要每天重複告警,只通知「新異常」與「恢復」
Day 22 還留下一個很重要的問題。
假設:
SERVER03
Disk = Critical
今天寄:
Critical Alert
明天還沒修:
又寄一次
後天:
又寄一次
久了就變成:
告警很多,但沒有人再看。
所以 Day 23 我們會第一次加入:
Previous State
例如保存:
SERVER01 = Healthy
SERVER02 = Warning
SERVER03 = Critical
下一次巡檢後比較:
Previous
vs
Current
判斷:
Healthy → Critical
= New Alert
Warning → Critical
= Escalated
Critical → Critical
= Existing Issue
Critical → Healthy
= Recovery
Healthy → Healthy
= No Notification
最後讓通知邏輯從:
現在有問題嗎?
升級成:
「這次跟上次相比,到底發生了什麼改變?」
Day 23 會正式讓我們的 Script 開始有 State,也會更接近真正 Monitoring System 的事件處理方式。