前面 28 天,我們其實已經累積了非常多零件。
從一開始最簡單的:
Get-Service
Get-Process
一路做到:
Server Health Check
AD User / Computer / Group Audit
Onboarding / Offboarding
PowerShell Remoting
CSV Configuration
Error Handling
Logging
Task Scheduler
HTML Report
Email Notification
State Tracking
PowerShell Module
Pester
Git
CI
Release
Deployment
Rollback
但現在如果打開專案,可能還是會看到:
這支 Script 做 Server
那支 Script 做 AD
另外一支做 HTML
再一支做 Email
State 又在另一支
每一個功能都會用。
但是:
它們還需要一個真正把所有東西串起來的入口。
所以 Day 29 不再加入新的大技術。
今天要做的是:
End-to-End Integration
把前 28 天做過的東西正式接成:
Task Scheduler
↓
DailyAutomation.ps1
↓
Load Configuration
↓
Import SysAdminToolkit
↓
Validate Environment
↓
Collect Server Health
↓
Collect AD Audit
↓
Analyze Current State
↓
Compare Previous State
↓
Generate Events
↓
Export Raw Data
↓
Generate HTML Report
↓
Notification
↓
Commit State
↓
Pipeline Summary
↓
Exit Code
到了今天,我們不再只是問:
這支 PowerShell Script 能不能跑?
而開始問:
整個 Automation Workflow 某一步失敗時,其他步驟應該怎麼處理?
今天最重要的第一個觀念:Health Status 跟 Execution Status 不一樣
這是做到 End-to-End Automation 後非常重要的一個區分。
假設:
SERVER01
CPU = 95%
OverallStatus = Critical
PowerShell:
成功連線
成功取得 CPU
成功判斷 Critical
成功產生 Report
成功發出通知
那這次 Automation 到底:
Success
還是:
Failed
Critical
Success
因為:
系統有問題,不代表 Automation 執行失敗。
相反地,Automation 正確偵測出 Critical,還把通知送出去,代表 Automation 本身其實成功完成工作。
所以 Day 29 開始分成兩種 Status
Infrastructure Health Status
描述:
系統現在健不健康?
例如:
Healthy
Warning
Critical
Unknown
Automation Execution Status
描述:
自動化流程自己有沒有正常完成?
例如:
Success
Partial
Failed
這兩個不要再混在一起。
一個非常典型的例子
假設:
SERVER01
CPU 95%
Critical
結果:
Collection Success
Analysis Success
HTML Success
Notification Success
State Commit Success
Critical
Success
這代表:
系統真的有問題,但我們的 Automation 正常工作。
另一種:
Servers 全部 Healthy
但是:
HTML Report
Failed
Healthy
Partial
因為:
基礎設施可能沒問題,但是 Automation Workflow 有一部分沒完成。
這個區分會讓後面的 Log、Dashboard、Task Scheduler 與 Troubleshooting 清楚很多。
Day 20 的 Exit Code 可以再進化
Day 20 我們曾經設計:
0
Success
1
Failed
2
Partial
今天保留這套設計。
但意義更明確:
exit 0
→ Automation Workflow 正常完成
exit 1
→ Fatal Automation Failure
exit 2
→ Automation 有完成主要工作,但部分 Stage 失敗
注意:
Critical Server
本身不一定造成 exit 1 或 exit 2。
它應該由:
HealthStatus
以及:
Notification
處理。
如果 Server 完全連不到呢?
以前我們可能:
SERVER02 WinRM Failed
→ Automation Partial
到了 Day 29 可以做得更細。
如果:
Get-ServerHealth
Failed
Unknown
並且成功回傳 Result Object:
[PSCustomObject]@{
ComputerName = "SERVER02"
Connection = "Failed"
OverallStatus = "Unknown"
ErrorMessage = "WinRM connection failed"
}
那其實:
Collector 自己是正常工作的。
它成功告訴我們:
「SERVER02 現在看不到。」
Unknown
Success
這比把:
Target Server 掛掉
跟:
PowerShell Script 自己壞掉
混在一起更成熟。
什麼才真的算 Automation Partial?
例如:
Server Collection
→ Success
AD Audit
→ Failed
HTML Report
→ Success
Notification
→ Success
Partial
因為 AD 資料沒有完成。
或者:
Server Collection
→ Success
HTML Report
→ Success
Critical Detected
→ Success
Email Notification
→ Failed
Partial
因為雖然問題有偵測到:
人沒有被成功通知。
什麼叫 Fatal Failure?
例如:
Production.psd1
讀不到
或:
SysAdminToolkit
無法 Import
或:
Config 格式完全錯誤
Failed
exit 1
今天先整理最終專案架構
到 Day 29:
SysAdmin-Automation/
│
├── Modules/
│ │
│ └── SysAdminToolkit/
│ ├── SysAdminToolkit.psd1
│ ├── SysAdminToolkit.psm1
│ │
│ ├── Public/
│ │ ├── Write-Log.ps1
│ │ ├── Get-ServerHealth.ps1
│ │ ├── Get-ADInactiveUserAudit.ps1
│ │ ├── Get-ADInactiveComputerAudit.ps1
│ │ ├── Get-ADGroupAudit.ps1
│ │ ├── Get-HealthEventType.ps1
│ │ ├── Save-HealthState.ps1
│ │ ├── New-DailyInfrastructureReport.ps1
│ │ └── Send-InfrastructureNotification.ps1
│ │
│ └── Private/
│
├── Config/
│ ├── Production.psd1
│ ├── Servers.csv
│ └── NotificationConfig.example.csv
│
├── Input/
│
├── Scripts/
│ ├── DailyAutomation.ps1
│ ├── Invoke-CITests.ps1
│ ├── Build-Release.ps1
│ └── Install-SysAdminToolkit.ps1
│
├── Tests/
│
├── Reports/
│
├── Logs/
│
├── State/
│
├── Artifacts/
│
├── README.md
├── CHANGELOG.md
└── Jenkinsfile
今天真正的主角就是:
Scripts/
└── DailyAutomation.ps1
為什麼 Main Script 不要再塞 1,500 行?
Day 24 我們已經把:
How
放進 Module。
例如:
CPU 怎麼查?
Memory 怎麼算?
Log 怎麼寫?
State 怎麼比較?
HTML 怎麼產?
交給:
SysAdminToolkit
所以:
DailyAutomation.ps1
應該主要描述:
Workflow
也就是:
先做什麼
成功後做什麼
失敗後怎麼處理
最後怎麼判斷整體結果
Main Script 應該讓工程師打開之後很快看懂:
今天整套 Automation 是怎麼跑的。
先建立 Production Configuration
例如:
Config/
└── Production.psd1
內容:
@{
Environment =
'Production'
ToolkitVersion =
'1.2.0'
Paths = @{
Servers =
'Config\Servers.csv'
Reports =
'Reports'
Logs =
'Logs'
State =
'State\ServerHealthState.json'
}
AD = @{
Enabled =
$true
UserSearchBase =
'OU=Users,DC=contoso,DC=com'
ComputerSearchBase =
'OU=Computers,DC=contoso,DC=com'
InactiveDays =
180
NeverLoggedOnGraceDays =
30
}
Notification = @{
Enabled =
$true
From =
'automation@contoso.com'
To =
'it-ops@contoso.com'
SmtpServer =
'smtp-relay.contoso.com'
}
}
注意這裡仍然沒有:
Password
Token
Credential
Secret 不應該直接放進:
Production.psd1
Module Version 也直接 Pin
例如:
ToolkitVersion =
'1.2.0'
主程式:
Import-Module SysAdminToolkit
-RequiredVersion $Config.ToolkitVersion
-ErrorAction Stop
這就接回 Day 28:
Production 應該明確知道自己正在執行哪個版本。
Servers.csv 繼續保留
例如:
ComputerName,Role,Environment,CPUWarning,CPUCritical,MemoryWarning,MemoryCritical
APP01,Application,Production,80,90,80,90
DB01,Database,Production,75,90,80,90
FILE01,FileServer,Production,80,90,80,90
也就是:
Production.psd1
→ Global Configuration
Servers.csv
→ Server-specific Configuration
每一次 Pipeline 都給一個 RunId
這是一個很實用的做法。
例如:
$RunId =
Get-Date `
-Format "yyyyMMdd_HHmmss"
可能:
20261007_060000
這一輪所有東西:
Log
CSV
HTML
State Event
Pipeline Summary
都可以用同一個 RunId。
例如:
Logs/
DailyAutomation_20261007_060000.log
Reports/
ServerHealth_20261007_060000.csv
ADAuditUsers_20261007_060000.csv
StateEvents_20261007_060000.csv
Pipeline_20261007_060000.csv
DailyReport_20261007_060000.html
以後看到某個 Report:
可以直接找到同一輪的 Log。
Bootstrap 階段有一個小問題
我們的:
Write-Log
在:
SysAdminToolkit
裡。
但是如果:
Import-Module
本身失敗,
那:
Write-Log
根本還不存在。
所以 Main Script 最前面最好保留一個非常簡單的:
Bootstrap Logger
只處理 Module 載入前的 Error。
Bootstrap Log
例如:
function Write-BootstrapLog {
param (
[string]$Path,
[string]$Message
)
$Time =
Get-Date `
-Format "yyyy-MM-dd HH:mm:ss"
Add-Content `
-Path $Path `
-Value "$Time $Message" `
-Encoding UTF8
}
它不用很複雜。
因為真正 Module 載入成功後:
Write-Log
會接手。
Pipeline Stage 也應該有自己的 Result
今天不只記:
Script Success
而是記每個 Stage。
例如:
Configuration
Module
ServerCollection
ADAudit
StateAnalysis
RawReport
HtmlReport
Notification
StateCommit
每一個都可以:
Success
Partial
Failed
Skipped
NotRequired
建立 Stage Result Object
例如:
function New-StageResult {
param (
[string]$RunId,
[string]$Stage,
[string]$Status,
[string]$Message,
[datetime]$StartTime
)
$EndTime =
Get-Date
[PSCustomObject]@{
RunId =
$RunId
Stage =
$Stage
Status =
$Status
StartTime =
$StartTime
EndTime =
$EndTime
DurationSeconds =
[math]::Round(
(
$EndTime -
$StartTime
).TotalSeconds,
2
)
Message =
$Message
}
}
最後就可以得到:
Stage Status Duration
Configuration Success 0.12
Module Success 0.30
ServerCollection Success 18.21
ADAudit Success 3.45
HtmlReport Success 0.48
Notification Success 1.25
StateCommit Success 0.06
這會讓:
Automation 自己也變得可以被觀察。
Stage Results
先:
$StageResults = @()
每完成一段:
$StageResults +=
New-StageResult ...
最後輸出:
Pipeline_20261007_060000.csv
Step 1:Load Configuration
首先:
$StageStart =
Get-Date
然後:
try {
$Config =
Import-PowerShellDataFile `
-Path $ConfigPath `
-ErrorAction Stop
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "Configuration" `
-Status "Success" `
-Message "Configuration loaded." `
-StartTime $StageStart
}
catch {
Write-BootstrapLog `
-Path $BootstrapLog `
-Message "Configuration failed: $($_.Exception.Message)"
exit 1
}
Config 完全讀不到:
後面沒有繼續的意義。
所以:
Fatal
Step 2:Import SysAdminToolkit
$StageStart =
Get-Date
try {
Import-Module `
SysAdminToolkit `
-RequiredVersion `
$Config.ToolkitVersion `
-Force `
-ErrorAction Stop
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "Module" `
-Status "Success" `
-Message "SysAdminToolkit $($Config.ToolkitVersion) loaded." `
-StartTime $StageStart
}
catch {
Write-BootstrapLog `
-Path $BootstrapLog `
-Message "Module import failed: $($_.Exception.Message)"
exit 1
}
這同樣是:
Fatal
Module 載入成功後開始正式 Write-Log
例如:
Write-Log -Path $LogFile
-Message "Daily Automation started."
Write-Log -Path $LogFile
-Message "RunId: $RunId"
Write-Log -Path $LogFile
-Message "Environment: $($Config.Environment)"
Write-Log -Path $LogFile
-Message "Toolkit Version: $($Config.ToolkitVersion)"
Log:
2026-10-07 06:00:00 [INFO] Daily Automation started.
2026-10-07 06:00:00 [INFO] RunId: 20261007_060000
2026-10-07 06:00:00 [INFO] Environment: Production
2026-10-07 06:00:00 [INFO] Toolkit Version: 1.2.0
Step 3:Collect Server Health
先讀:
$Servers =
Import-Csv `
$ServersCsv
然後:
$ServerResults =
foreach ($Server in $Servers) {
Write-Log `
-Path $LogFile `
-Message "Checking $($Server.ComputerName)"
Get-ServerHealth `
-ComputerName `
$Server.ComputerName `
-CPUWarning `
([int]$Server.CPUWarning) `
-CPUCritical `
([int]$Server.CPUCritical) `
-MemoryWarning `
([int]$Server.MemoryWarning) `
-MemoryCritical `
([int]$Server.MemoryCritical)
}
注意:
SERVER01 Connection Failed
仍然應該回:
ComputerName = SERVER01
Connection = Failed
OverallStatus = Unknown
而不是讓 SERVER01:
從結果中消失。
判斷整體 Server Health
$CriticalCount =
@(
$ServerResults |
Where-Object {
$_.OverallStatus -eq "Critical"
}
).Count
$WarningCount =
@(
$ServerResults |
Where-Object {
$_.OverallStatus -eq "Warning"
}
).Count
$UnknownCount =
@(
$ServerResults |
Where-Object {
$_.OverallStatus -eq "Unknown"
}
).Count
然後:
if ($CriticalCount -gt 0) {
$HealthStatus =
"Critical"
}
elseif ($UnknownCount -gt 0) {
$HealthStatus =
"Unknown"
}
elseif ($WarningCount -gt 0) {
$HealthStatus =
"Warning"
}
else {
$HealthStatus =
"Healthy"
}
為什麼 Unknown 放在 Warning 前面?
因為:
Warning
至少代表:
我知道現在發生什麼。
但是:
Unknown
代表:
我失去了 Visibility。
所以在這個簡化架構裡:
Critical
Unknown
Warning
Healthy
是合理的顯示順序。
但這不是全球唯一標準,正式環境可以依公司的 Monitoring Policy 調整。
Step 4:AD Audit
前面 Day 14~16 已經寫過:
Inactive Users
Inactive Computers
Group Audit
Day 29 不再重新貼一次全部查詢。
假設現在這些邏輯已整理成:
Get-ADInactiveUserAudit
Get-ADInactiveComputerAudit
Get-ADGroupAudit
主程式只要:
$InactiveUsers =
Get-ADInactiveUserAudit -SearchBase
$Config.AD.UserSearchBase -InactiveDays
$Config.AD.InactiveDays -NeverLoggedOnGraceDays
$Config.AD.NeverLoggedOnGraceDays
$InactiveComputers =
Get-ADInactiveComputerAudit -SearchBase
$Config.AD.ComputerSearchBase -InactiveDays
$Config.AD.InactiveDays
$GroupAudit =
Get-ADGroupAudit
這就是 Module 化的價值。
Main Script 不需要再知道:
lastLogonTimestamp 怎麼轉?
MemberOf 怎麼查?
Group Recursive 怎麼判斷?
它只知道:
我要執行 AD Audit。
如果 AD Audit 失敗
例如:
ActiveDirectory Module
有問題
但是 Server Health 已經成功取得。
這時候我不會:
exit 1
而是:
AD Stage = Failed
HasPartialError = True
然後:
其他工作繼續。
因為我們還是可以產生一份:
Server Health Report
只是 AD 區塊顯示:
No Data / Audit Failed
這就是 Graceful Degradation
正常:
Server
✓
AD
✓
HTML
✓
Email
✓
如果 AD 壞掉:
Server
✓
AD
✗
HTML
✓
但標示 AD No Data
Email
✓
並提醒 AD Audit Failed
而不是:
AD 出錯
↓
整個 Script 直接停止
↓
連 Server Report 都沒有
Step 5:Export Raw Data
不管 HTML 最後能不能成功,
Raw Data 最好先保存。
例如:
$ServerResults |
Export-Csv -Path $ServerReport
-NoTypeInformation `
-Encoding UTF8
AD:
$InactiveUsers |
Export-Csv -Path $InactiveUserReport
-NoTypeInformation `
-Encoding UTF8
Computer:
$InactiveComputers |
Export-Csv -Path $InactiveComputerReport
-NoTypeInformation `
-Encoding UTF8
這樣即使:
HTML Generator
失敗,
至少:
Raw Data 還在。
Step 6:State Comparison
接回 Day 23。
先讀:
State/
ServerHealthState.json
建立:
$PreviousState
再建立:
$CurrentState =
$ServerResults |
Select-Object `
ComputerName,
OverallStatus,
@{
Name = "CheckTime"
Expression = {
Get-Date `
-Format "yyyy-MM-dd HH:mm:ss"
}
}
建立 Previous Map
$PreviousMap =
@{}
然後:
foreach ($Item in $PreviousState) {
$PreviousMap[
$Item.ComputerName
] = $Item
}
Compare
$StateEvents =
foreach ($Current in $CurrentState) {
$PreviousStatus =
$null
if (
$PreviousMap.ContainsKey(
$Current.ComputerName
)
) {
$PreviousStatus =
$PreviousMap[
$Current.ComputerName
].OverallStatus
}
$EventType =
Get-HealthEventType `
-PreviousStatus `
$PreviousStatus `
-CurrentStatus `
$Current.OverallStatus
[PSCustomObject]@{
ComputerName =
$Current.ComputerName
PreviousStatus =
if ($PreviousStatus) {
$PreviousStatus
}
else {
"None"
}
CurrentStatus =
$Current.OverallStatus
EventType =
$EventType
CheckTime =
$Current.CheckTime
}
}
Notify Events
$NotifyEventTypes =
@(
"InitialIssue"
"NewAlert"
"Escalated"
"Recovery"
"VisibilityLost"
"StateRestoredWithIssue"
)
然後:
$NotifyEvents =
@(
$StateEvents |
Where-Object {
$_.EventType -in
$NotifyEventTypes
}
)
這樣通知就不再根據「現在有沒有 Critical」
而是根據:
這一次發生了什麼改變。
例如:
SERVER01
Healthy
→ Critical
NewAlert
寄信。
SERVER02
Critical
→ Critical
ExistingIssue
不重複寄。
SERVER03
Critical
→ Healthy
Recovery
寄 Recovery。
Step 7:HTML Report
現在 HTML Report 可以接收:
ServerResults
InactiveUsers
InactiveComputers
GroupAudit
StateEvents
HealthStatus
RunId
例如:
$HtmlReport =
New-DailyInfrastructureReport -RunId
$RunId -ServerResults
$ServerResults -InactiveUsers
$InactiveUsers -InactiveComputers
$InactiveComputers -GroupAudit
$GroupAudit -StateEvents
$StateEvents -OutputPath
$HtmlReportPath
這就是 Day 21 的 HTML Logic 搬進 Module 後的樣子。
HTML 第一眼應該看到
Daily Infrastructure Report
RunId
20261007_060000
Environment
Production
Infrastructure Health
Critical
Server Summary
Total 20
Healthy 16
Warning 2
Critical 1
Unknown 1
State Changes
APP01
Healthy → Critical
NewAlert
DB01
Unknown → Healthy
Recovery
AD Audit
Inactive Users 15
Inactive Computers 28
Group Review 3
Report 自己也要顯示 Data Source Status
例如:
Server Collection
Available
AD User Audit
Available
AD Computer Audit
Failed
不要讓:
AD Computer Count = 0
造成誤會。
因為:
真的 0 台 stale computer
跟:
AD Computer Query 根本沒成功
Query 成功
真的沒有資料
Query 沒有成功
Dashboard 必須分得出來。
Step 8:Notification
現在:
$NotifyEvents.Count
如果:
0
NotRequired
不需要寄信。
如果:
0
而且:
$Config.Notification.Enabled
為:
True
才:
Send-InfrastructureNotification
例如:
Send-InfrastructureNotification -Events
$NotifyEvents -HealthStatus
$HealthStatus -ReportPath
$HtmlReportPath -From
$Config.Notification.From -To
$Config.Notification.To -SmtpServer
$Config.Notification.SmtpServer `
-ErrorAction Stop
Notification Failure 仍然不能讓 Raw Data 消失
假設:
Collection
✓
CSV
✓
HTML
✓
State Analysis
✓
Email
✗
我們仍然保留:
CSV
HTML
Log
Execution:
Partial
而不是把前面的成果全部視為不存在。
Step 9:什麼時候 Commit State?
Day 23 已經提過這個坑。
假設:
Healthy
→ Critical
我們先把 State 更新成:
Critical
然後:
Email Failure
下一輪:
Critical
→ Critical
就會:
ExistingIssue
不再嘗試通知。
所以今天先繼續使用 Day 23 的簡化策略:
Generate Event
↓
Need Notification?
↓
Notify
↓
成功
↓
Commit Current State
如果不需要通知
例如:
Healthy
→ Healthy
可以直接:
Commit State
如果通知成功
NewAlert
↓
Email Success
↓
Commit State
如果通知失敗
NewAlert
↓
Email Failed
↓
Do Not Commit
保留 Previous State。
下次重新:
Healthy → Critical
再嘗試一次。
這是比較保守的設計
代價是:
一封 Email 失敗,可能讓其他同一輪 State 也暫時不 Commit。
更成熟的平台可以把:
Monitoring State
跟:
Notification Delivery State
拆成兩套資料。
例如:
CurrentHealthState.json
NotificationQueue.json
但 Day 29 的目標是:
把目前 29 天內容真正串起來。
所以先保留我們 Day 23 已經建立的簡化策略。
Step 10:Pipeline Summary
最後不要只有:
Script completed.
而是產生:
Pipeline Summary
例如:
$PipelineSummary =
[PSCustomObject]@{
RunId =
$RunId
Environment =
$Config.Environment
ToolkitVersion =
$Config.ToolkitVersion
HealthStatus =
$HealthStatus
ExecutionStatus =
$ExecutionStatus
TotalServers =
@($ServerResults).Count
CriticalServers =
$CriticalCount
WarningServers =
$WarningCount
UnknownServers =
$UnknownCount
NotifyEvents =
@($NotifyEvents).Count
NotificationStatus =
$NotificationStatus
StartTime =
$PipelineStart
EndTime =
Get-Date
}
輸出:
PipelineSummary_20261007_060000.csv
最重要的結果可能是這樣
RunId
20261007_060000
HealthStatus
Critical
ExecutionStatus
Success
TotalServers
20
CriticalServers
1
WarningServers
2
UnknownServers
0
NotifyEvents
1
NotificationStatus
Success
這個結果非常合理。
因為:
Infrastructure
有 Critical
但:
Automation
完整成功
另一種
HealthStatus
Warning
ExecutionStatus
Partial
NotificationStatus
Failed
代表:
系統目前有 Warning
+
Automation 也有部分失敗
這兩件事情同時存在。
End-to-End Main Script
下面整理一份比較完整的:
DailyAutomation.ps1
範例。
前面 Day 14~23 的細部 Function 不再全部重貼;假設它們已經整理進:
SysAdminToolkit
今天專注:
Orchestration。
param (
[string]
$ConfigPath
)
$ErrorActionPreference =
"Stop"
$PipelineStart =
Get-Date
$RunId =
Get-Date `
-Format "yyyyMMdd_HHmmss"
$ProjectRoot =
Split-Path -Path $PSScriptRoot
-Parent
if (
[string]::IsNullOrWhiteSpace(
$ConfigPath
)
) {
$ConfigPath =
Join-Path `
$ProjectRoot `
"Config\Production.psd1"
}
$BootstrapLogFolder =
Join-Path $ProjectRoot
"Logs"
if (-not (Test-Path $BootstrapLogFolder)) {
New-Item `
-Path $BootstrapLogFolder `
-ItemType Directory `
-Force |
Out-Null
}
$BootstrapLog =
Join-Path $BootstrapLogFolder
"DailyAutomation_$RunId.log"
function Write-BootstrapLog {
param (
[string]
$Path,
[string]
$Message
)
$Time =
Get-Date `
-Format "yyyy-MM-dd HH:mm:ss"
Add-Content `
-Path $Path `
-Value "$Time $Message" `
-Encoding UTF8
}
function New-StageResult {
param (
[string]
$RunId,
[string]
$Stage,
[string]
$Status,
[string]
$Message,
[datetime]
$StartTime
)
$EndTime =
Get-Date
[PSCustomObject]@{
RunId =
$RunId
Stage =
$Stage
Status =
$Status
StartTime =
$StartTime
EndTime =
$EndTime
DurationSeconds =
[math]::Round(
(
$EndTime -
$StartTime
).TotalSeconds,
2
)
Message =
$Message
}
}
$StageResults =
@()
$HasPartialError =
$false
$StageStart =
Get-Date
try {
$Config =
Import-PowerShellDataFile `
-Path $ConfigPath `
-ErrorAction Stop
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "Configuration" `
-Status "Success" `
-Message "Configuration loaded." `
-StartTime $StageStart
}
catch {
Write-BootstrapLog `
-Path $BootstrapLog `
-Message "Configuration failed: $($_.Exception.Message)"
exit 1
}
$StageStart =
Get-Date
try {
Import-Module `
SysAdminToolkit `
-RequiredVersion `
$Config.ToolkitVersion `
-Force `
-ErrorAction Stop
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "Module" `
-Status "Success" `
-Message "SysAdminToolkit $($Config.ToolkitVersion) loaded." `
-StartTime $StageStart
}
catch {
Write-BootstrapLog `
-Path $BootstrapLog `
-Message "Module import failed: $($_.Exception.Message)"
exit 1
}
$ServersCsv =
Join-Path $ProjectRoot
$Config.Paths.Servers
$ReportFolder =
Join-Path $ProjectRoot
$Config.Paths.Reports
$LogFolder =
Join-Path $ProjectRoot
$Config.Paths.Logs
$StateFile =
Join-Path $ProjectRoot
$Config.Paths.State
foreach ($Folder in @(
$ReportFolder,
$LogFolder
)) {
if (-not (Test-Path $Folder)) {
New-Item `
-Path $Folder `
-ItemType Directory `
-Force |
Out-Null
}
}
$StateFolder =
Split-Path $StateFile
-Parent
if (-not (Test-Path $StateFolder)) {
New-Item `
-Path $StateFolder `
-ItemType Directory `
-Force |
Out-Null
}
$LogFile =
Join-Path $LogFolder
"DailyAutomation_$RunId.log"
Write-Log -Path $LogFile
-Message "Daily Automation started."
Write-Log -Path $LogFile
-Message "RunId: $RunId"
Write-Log -Path $LogFile
-Message "Toolkit Version: $($Config.ToolkitVersion)"
$StageStart =
Get-Date
try {
$Servers =
@(
Import-Csv `
-Path $ServersCsv `
-ErrorAction Stop
)
if ($Servers.Count -eq 0) {
throw `
"Server configuration is empty."
}
$ServerResults =
foreach ($Server in $Servers) {
Write-Log `
-Path $LogFile `
-Message "Checking $($Server.ComputerName)"
Get-ServerHealth `
-ComputerName `
$Server.ComputerName `
-CPUWarning `
([int]$Server.CPUWarning) `
-CPUCritical `
([int]$Server.CPUCritical) `
-MemoryWarning `
([int]$Server.MemoryWarning) `
-MemoryCritical `
([int]$Server.MemoryCritical)
}
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "ServerCollection" `
-Status "Success" `
-Message "$($ServerResults.Count) server results collected." `
-StartTime $StageStart
}
catch {
$ServerResults =
@()
$HasPartialError =
$true
Write-Log `
-Path $LogFile `
-Message "Server collection failed: $($_.Exception.Message)" `
-Level "ERROR"
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "ServerCollection" `
-Status "Failed" `
-Message $_.Exception.Message `
-StartTime $StageStart
}
$CriticalCount =
@(
$ServerResults |
Where-Object {
$_.OverallStatus -eq "Critical"
}
).Count
$WarningCount =
@(
$ServerResults |
Where-Object {
$_.OverallStatus -eq "Warning"
}
).Count
$UnknownCount =
@(
$ServerResults |
Where-Object {
$_.OverallStatus -eq "Unknown"
}
).Count
if ($CriticalCount -gt 0) {
$HealthStatus =
"Critical"
}
elseif ($UnknownCount -gt 0) {
$HealthStatus =
"Unknown"
}
elseif ($WarningCount -gt 0) {
$HealthStatus =
"Warning"
}
elseif ($ServerResults.Count -gt 0) {
$HealthStatus =
"Healthy"
}
else {
$HealthStatus =
"Unknown"
}
$StageStart =
Get-Date
$InactiveUsers =
@()
$InactiveComputers =
@()
$GroupAudit =
@()
if ($Config.AD.Enabled) {
try {
$InactiveUsers =
@(
Get-ADInactiveUserAudit `
-SearchBase `
$Config.AD.UserSearchBase `
-InactiveDays `
$Config.AD.InactiveDays `
-NeverLoggedOnGraceDays `
$Config.AD.NeverLoggedOnGraceDays
)
$InactiveComputers =
@(
Get-ADInactiveComputerAudit `
-SearchBase `
$Config.AD.ComputerSearchBase `
-InactiveDays `
$Config.AD.InactiveDays
)
$GroupAudit =
@(
Get-ADGroupAudit
)
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "ADAudit" `
-Status "Success" `
-Message "AD audit completed." `
-StartTime $StageStart
}
catch {
$HasPartialError =
$true
Write-Log `
-Path $LogFile `
-Message "AD audit failed: $($_.Exception.Message)" `
-Level "ERROR"
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "ADAudit" `
-Status "Failed" `
-Message $_.Exception.Message `
-StartTime $StageStart
}
}
else {
$StageResults +=
New-StageResult `
-RunId $RunId `
-Stage "ADAudit" `
-Status "Skipped" `
-Message "AD audit disabled." `
-StartTime $StageStart
}
$ServerReport =
Join-Path $ReportFolder
"ServerHealth_$RunId.csv"
$InactiveUserReport =
Join-Path $ReportFolder
"InactiveUsers_$RunId.csv"
$InactiveComputerReport =
Join-Path $ReportFolder
"InactiveComputers_$RunId.csv"
$ServerResults |
Export-Csv -Path $ServerReport
-NoTypeInformation `
-Encoding UTF8
$InactiveUsers |
Export-Csv -Path $InactiveUserReport
-NoTypeInformation `
-Encoding UTF8
$InactiveComputers |
Export-Csv -Path $InactiveComputerReport
-NoTypeInformation `
-Encoding UTF8
try {
if (Test-Path $StateFile) {
$PreviousState =
@(
Get-Content `
-Path $StateFile `
-Raw `
-ErrorAction Stop |
ConvertFrom-Json `
-ErrorAction Stop
)
}
else {
$PreviousState =
@()
}
}
catch {
$HasPartialError =
$true
$PreviousState =
@()
Write-Log `
-Path $LogFile `
-Message "State load failed: $($_.Exception.Message)" `
-Level "ERROR"
}
$CurrentState =
@(
$ServerResults |
Select-Object `
ComputerName,
OverallStatus,
@{
Name =
"CheckTime"
Expression = {
Get-Date `
-Format "yyyy-MM-dd HH:mm:ss"
}
}
)
$PreviousMap =
@{}
foreach ($Item in $PreviousState) {
$PreviousMap[
$Item.ComputerName
] = $Item
}
$StateEvents =
foreach ($Current in $CurrentState) {
$PreviousStatus =
$null
if (
$PreviousMap.ContainsKey(
$Current.ComputerName
)
) {
$PreviousStatus =
$PreviousMap[
$Current.ComputerName
].OverallStatus
}
$EventType =
Get-HealthEventType `
-PreviousStatus `
$PreviousStatus `
-CurrentStatus `
$Current.OverallStatus
[PSCustomObject]@{
ComputerName =
$Current.ComputerName
PreviousStatus =
if ($PreviousStatus) {
$PreviousStatus
}
else {
"None"
}
CurrentStatus =
$Current.OverallStatus
EventType =
$EventType
CheckTime =
$Current.CheckTime
}
}
$StateEventReport =
Join-Path $ReportFolder
"StateEvents_$RunId.csv"
$StateEvents |
Export-Csv -Path $StateEventReport
-NoTypeInformation `
-Encoding UTF8
$NotifyEventTypes =
@(
"InitialIssue"
"NewAlert"
"Escalated"
"Recovery"
"VisibilityLost"
"StateRestoredWithIssue"
)
$NotifyEvents =
@(
$StateEvents |
Where-Object {
$_.EventType -in
$NotifyEventTypes
}
)
$HtmlReportPath =
Join-Path $ReportFolder
"DailyReport_$RunId.html"
try {
New-DailyInfrastructureReport `
-RunId $RunId `
-HealthStatus $HealthStatus `
-ServerResults $ServerResults `
-InactiveUsers $InactiveUsers `
-InactiveComputers $InactiveComputers `
-GroupAudit $GroupAudit `
-StateEvents $StateEvents `
-OutputPath $HtmlReportPath
$HtmlStatus =
"Success"
}
catch {
$HasPartialError =
$true
$HtmlStatus =
"Failed"
Write-Log `
-Path $LogFile `
-Message "HTML report failed: $($_.Exception.Message)" `
-Level "ERROR"
}
$NotificationStatus =
"NotRequired"
$CanCommitState =
$true
if (
$NotifyEvents.Count -gt 0 -and
$Config.Notification.Enabled
) {
try {
Send-InfrastructureNotification `
-Events $NotifyEvents `
-HealthStatus $HealthStatus `
-ReportPath $HtmlReportPath `
-From $Config.Notification.From `
-To $Config.Notification.To `
-SmtpServer $Config.Notification.SmtpServer `
-ErrorAction Stop
$NotificationStatus =
"Success"
$CanCommitState =
$true
}
catch {
$NotificationStatus =
"Failed"
$CanCommitState =
$false
$HasPartialError =
$true
Write-Log `
-Path $LogFile `
-Message "Notification failed: $($_.Exception.Message)" `
-Level "ERROR"
}
}
elseif (
$NotifyEvents.Count -gt 0
) {
$NotificationStatus =
"Disabled"
$CanCommitState =
$true
}
if ($CanCommitState) {
try {
Save-HealthState `
-State $CurrentState `
-Path $StateFile
$StateCommitStatus =
"Success"
}
catch {
$StateCommitStatus =
"Failed"
$HasPartialError =
$true
Write-Log `
-Path $LogFile `
-Message "State commit failed: $($_.Exception.Message)" `
-Level "ERROR"
}
}
else {
$StateCommitStatus =
"Skipped"
Write-Log `
-Path $LogFile `
-Message "State commit skipped because notification failed." `
-Level "WARNING"
}
if ($HasPartialError) {
$ExecutionStatus =
"Partial"
}
else {
$ExecutionStatus =
"Success"
}
$PipelineSummary =
[PSCustomObject]@{
RunId =
$RunId
Environment =
$Config.Environment
ToolkitVersion =
$Config.ToolkitVersion
HealthStatus =
$HealthStatus
ExecutionStatus =
$ExecutionStatus
TotalServers =
@($ServerResults).Count
CriticalServers =
$CriticalCount
WarningServers =
$WarningCount
UnknownServers =
$UnknownCount
InactiveUsers =
@($InactiveUsers).Count
InactiveComputers =
@($InactiveComputers).Count
NotifyEvents =
@($NotifyEvents).Count
HtmlStatus =
$HtmlStatus
NotificationStatus =
$NotificationStatus
StateCommitStatus =
$StateCommitStatus
StartTime =
$PipelineStart
EndTime =
Get-Date
}
$PipelineSummaryPath =
Join-Path $ReportFolder
"PipelineSummary_$RunId.csv"
$PipelineSummary |
Export-Csv -Path $PipelineSummaryPath
-NoTypeInformation `
-Encoding UTF8
$StageReport =
Join-Path $ReportFolder
"PipelineStages_$RunId.csv"
$StageResults |
Export-Csv -Path $StageReport
-NoTypeInformation `
-Encoding UTF8
Write-Log -Path $LogFile
-Message "Health Status: $HealthStatus"
Write-Log -Path $LogFile
-Message "Execution Status: $ExecutionStatus"
Write-Log -Path $LogFile
-Message "Daily Automation completed."
if (
$ExecutionStatus -eq
"Partial"
) {
exit 2
}
exit 0
這支 Main Script 最大的重點不是 Code
真正重要的是它現在有清楚的責任分工:
Config
↓
Module
↓
Collection
↓
Analysis
↓
State
↓
Report
↓
Notification
↓
State Commit
↓
Summary
不像以前:
Function
Function
Function
foreach
if
try
CSV
SMTP
AD
全部混在一起
Task Scheduler 現在只需要啟動 Master Script
Day 20 設定:
Program
powershell.exe
Arguments:
-NoProfile
-NonInteractive
-File "C:\SysAdmin-Automation\Scripts\DailyAutomation.ps1"
-ConfigPath "C:\SysAdmin-Automation\Config\Production.psd1"
Task Scheduler 的責任只有:
準時啟動 Pipeline。
至於:
Server
AD
HTML
Notification
State
全部交給:
DailyAutomation.ps1
現在 Task Scheduler 看到 Exit Code
0
代表:
Automation Execution Success
不代表:
所有 Server Healthy
Critical
Success
Task Scheduler:
0
完全合理。
真正的 Critical 會透過:
Dashboard
+
Notification
處理。
如果 Exit Code = 2
代表:
Automation Partial
例如:
HTML Failed
或:
Notification Failed
或:
AD Audit Failed
這時才是:
Automation 自己需要被維護。
Automation 也需要被 Monitoring
做到 Day 29 會發現一件很有趣的事情。
我們原本寫 Automation 是為了監控:
Server
AD
但現在:
Automation 本身也可能壞。
例如:
Task Scheduler 沒跑
Module Import Failed
State 壞掉
HTML Generator Failed
SMTP Failed
所以成熟一點之後:
Monitoring System
甚至應該監控:
Automation Last Run Time
Automation Exit Code
Report 是否產生
State 是否更新
這就是:
Who watches the watcher?
我們今天先透過:
Pipeline Summary
Stage Results
Exit Code
Log
留下基礎。
一個早上的實際流程
假設每天:
06:00
Task Scheduler 啟動。
06:00:
Load Config
Success
06:00:
Load SysAdminToolkit 1.2.0
Success
06:01:
20 Servers
開始巡檢
結果:
17 Healthy
2 Warning
1 Critical
Health:
Critical
06:02:
AD Audit
找到:
15 Inactive Users
28 Inactive Computers
3 Group Reviews
06:03:
State 比較:
APP01
Healthy
→ Critical
NewAlert
另一台:
DB01
Critical
→ Critical
ExistingIssue
不重新通知。
另一台:
FILE01
Warning
→ Healthy
Recovery
通知 Recovery。
06:04:
CSV
Success
06:04:
HTML
Success
06:05:
Email
NewAlert
APP01
Recovery
FILE01
成功。
06:05:
Commit State
Success
Critical
Success
0
這就是一個真正合理的 Daily Automation Pipeline。
如果 SMTP 掛掉
同一輪:
Collection
Success
AD
Success
HTML
Success
NewAlert
Detected
Email
Failed
Critical
Partial
Failed
Skipped
2
下次執行仍然:
Healthy → Critical
再次嘗試通知。
如果 Config 壞掉
例如:
Production.psd1
Syntax Error
那:
Configuration
Failed
後面:
Server Collection
根本無法開始
這就是:
Fatal
1
這跟:
SERVER01 CPU Critical
完全是不同等級的事情。
如果 AD 壞掉,但 Server 正常
例如:
Server Collection
Success
AD Audit
Failed
HTML
Success
HTML 可以顯示:
Server Health
Healthy
AD Audit
No Data
Collection Failed
Partial
2
這就是:
Partial Failure
Day 19 我們第一次正式談這個概念。
到了 Day 29,它已經成為:
整個 Automation Pipeline 的核心設計原則。
不要因為一個 Stage 出錯就全部 Abort
需要 Abort 的:
Configuration 完全無法讀
Module 完全無法載入
核心環境無法初始化
可以 Continue 的:
AD Audit Failure
HTML Failure
Email Failure
某一個 Server Unknown
但 Continue 不代表忽略。
而是:
繼續可以完成的工作
+
把錯誤記下來
+
最後標示 Partial
這就是:
Resilient Automation。
Production Change Workflow 不應該直接塞進 Daily Pipeline
這一點也很重要。
Day 18、Day 19 我們做過:
Onboarding
Offboarding
但今天 Daily Pipeline 裡沒有直接:
06:00
↓
自動建立所有帳號
↓
自動停用所有離職者
因為:
Server Health Check
跟:
Disable User
風險完全不同。
Read-Only Automation
非常適合:
Scheduled
Fully Automated
例如:
Server Health
AD Audit
Event Log
Reports
State Compare
Notification
Change Automation
例如:
New-ADUser
Disable-ADAccount
Remove-ADGroupMember
仍然應該:
Request
↓
Validation
↓
Preview
↓
Approval
↓
Execution
↓
Verification
↓
Audit
不要因為我們現在有:
Task Scheduler
CI/CD
就認為:
所有 Change 都應該無條件全自動。
自動化的成熟不是「拿掉所有人」
真正成熟比較像:
Machine
負責
重複
大量
可預測
可驗證
人負責:
Risk
Approval
Business Context
Exception
Decision
這也是這 29 天一直在強調的:
Automation 不應該拿掉 Judgment,而是把 Judgment 放在真正需要它的地方。
CI Pipeline 跟 Daily Pipeline 也不要混在一起
目前我們已經有兩條 Pipeline。
Code Pipeline
Day 27~28:
Developer
↓
Git
↓
Pester
↓
CI
↓
Release
↓
Artifact
↓
Deploy
它管理:
Automation Code。
Operations Pipeline
今天:
Task Scheduler
↓
Collect
↓
Analyze
↓
Report
↓
Notify
↓
State
它管理:
每天的 IT Operations。
兩條 Pipeline 最後會接起來
Development Pipeline
│
▼
Operations Pipeline
Task Scheduler
↓
DailyAutomation.ps1
↓
SysAdminToolkit 1.2.0
↓
Server / AD
↓
State
↓
Report
↓
Notification
這其實就是這整個系列從 Day 1 一路走到 Day 29 的成果。
我們不再只有一支 PowerShell Script
現在其實已經有:
Code
+
Configuration
+
Module
+
Testing
+
Version Control
+
CI
+
Release
+
Deployment
+
Scheduling
+
Logging
+
State
+
Reporting
+
Notification
也就是:
Automation Platform 的雛形
雖然它仍然是:
PowerShell
+
Windows Server
+
Active Directory
但背後的工程概念已經開始跟:
DevOps
Platform Engineering
SRE
重疊。
Day 29 小結
Day 29 沒有再新增:
新的 PowerShell Cmdlet
而是把前面 28 天真正串起來。
我們建立:
DailyAutomation.ps1
當作整個日常維運 Automation 的:
Orchestrator。
現在完整流程是:
Task Scheduler
↓
Configuration
↓
SysAdminToolkit
↓
Server Collection
↓
AD Audit
↓
Current State
↓
Previous State
↓
State Comparison
↓
Event
↓
Raw CSV
↓
HTML Report
↓
Notification
↓
State Commit
↓
Pipeline Summary
↓
Exit Code
今天最重要的一個觀念,就是正式分開:
HealthStatus
與:
ExecutionStatus
Critical
Success
完全合理。
因為:
Automation 的工作就是正確地發現 Critical,而不是把所有 Critical 都當成 Script Failure。
我們也把 Day 19 的:
Partial Failure
正式提升到整個 Pipeline。
不是:
一個地方出錯
↓
全部停止
而是:
能做的繼續
↓
不能做的留下 Error
↓
最後標記 Partial
並產生:
Stage Results
Pipeline Summary
Log
Exit Code
讓:
Automation 本身也可以被 Troubleshoot。
現在這 29 天的成果,已經從 Day 1:
Get-Service
慢慢長成:
Git
│
▼
Pester
│
▼
CI
│
▼
Release
│
▼
Deploy
│
▼
SysAdminToolkit
│
▼
Task Scheduler
│
▼
DailyAutomation.ps1
│
┌────────────┼────────────┐
▼ ▼ ▼
Windows Server AD State
│ │ │
└────────────┼────────────┘
▼
Analyze
│
▼
CSV
│
▼
HTML
│
▼
Notification
│
▼
Engineer Review
這已經不只是:
「我會寫 PowerShell。」
而是開始變成:
「我可以設計、測試、部署並維護一套 PowerShell Automation Workflow。」
Day 30 預告|30 天完整總結
Day 30|從 PowerShell Script 到 SysAdmin Automation Toolkit:30 天完整回顧與最終架構
Day 30 我不會再塞新的技術。
會按照我們前面約定的方式,做整個系列的:
完整歸納與收尾
會從:
Day 1
為什麼需要 Automation?
一路重新串到:
Day 29
End-to-End Automation Pipeline
並整理成幾個完整階段:
PowerShell Foundation
↓
Server Automation
↓
Remote Operations
↓
Active Directory Automation
↓
Scheduling
↓
Reporting
↓
Notification
↓
State Tracking
↓
Module
↓
Testing
↓
Git
↓
CI
↓
Release / Deployment
↓
End-to-End Pipeline
Day 30 也會把目前散落在文章裡的所有東西整理成一個最終:
SysAdmin Automation Toolkit
包含最後的:
Modules/
Scripts/
Config/
Input/
Tests/
Reports/
Logs/
State/
Artifacts/
以及:
哪些工作可以全自動?
哪些 Change 一定要 Approval?
哪些資料要留下 Audit Trail?
什麼叫 Success / Partial / Failed?
Health Status 跟 Execution Status 怎麼分?