Day 20 我們完成了第一個完整 System Design Case Study:
Design URL Shortener
我們從:
Requirement
↓
Traffic Estimation
↓
API
↓
Data Model
↓
Cache
↓
Replication
↓
Sharding
↓
Reliability
↓
Security
一步一步把 Architecture 推導出來。
Day 21 我們來設計另一個非常常見,而且會大量使用前面所學觀念的系統:
Notification System(通知系統)
例如:
Amazon:
Your package has been shipped.
Bank:
Your credit card was charged $100.
Instagram:
Alvin liked your photo.
Google:
New sign-in detected.
Food Delivery:
Your order is arriving soon.
這些其實都是 Notification。
看起來只是:
Send Message
但當 System 需要每天發送:
Millions
甚至
Billions
的 Notifications,就會遇到很多問題:
Email、SMS、Push 有什麼不同?
Notification 要同步送還是非同步送?
Message Queue 為什麼重要?
Provider 掛掉怎麼辦?
Notification 發送失敗要不要 Retry?
Retry 會不會造成重複通知?
什麼是 Idempotency?
什麼是 Deduplication?
User 不想收到 Marketing Email 怎麼辦?
不同 Notification 有沒有不同 Priority?
Scheduled Notification 怎麼做?
如何知道 Notification 到底有沒有成功?
DLQ 是什麼?
APNs 是什麼?
FCM 是什麼?
大量 Notification 會不會把 Provider 打爆?
今天我們就從 0 開始。
而且每遇到新的縮寫,我們都會先拆開解釋。
Notification(通知):
System 主動把某個重要資訊告訴 User。
例如:
Order shipped
Payment successful
Password changed
New message received
Promotion available
Notification 和一般 API Response 不太一樣。
例如 User Checkout:
User
↓
POST /checkout
↓
Order Service
↓
Response:
Order created
這是 User 主動 Request,Server 回覆。
但 Notification 可能是:
Order created
↓
30 minutes later
↓
Package shipped
↓
System 主動通知 User
所以 Notification 常常是:
某個 Event 發生之後,System 主動把資訊送出去。
Event(事件):
描述「某件事情已經發生」的資料或訊息。
例如:
ORDER_CREATED
PAYMENT_SUCCESSFUL
PACKAGE_SHIPPED
PASSWORD_CHANGED
注意:
Event
不是「叫某個 Service 做事情」。
它比較像是在說:
某件事情已經發生了。
例如:
PACKAGE_SHIPPED
意思是:
包裹已經寄出了
其他 System 可以根據這個 Event 決定:
Send Email
Send Push Notification
Update Analytics
User 可以透過不同方式收到 Notification。
這些不同方式叫:
Channel(通道 / 通知管道)
在 Notification System 裡:
Channel = Notification 送到 User 的方式。
常見:
Email
SMS
Push Notification
In-App Notification
Email = Electronic Mail
中文:
電子郵件
例如:
Your order has shipped.
寄到:
user@example.com
適合:
Order Confirmation
Receipt
Password Reset
Newsletter
Security Alert
SMS = Short Message Service
拆開:
Short → 短的
Message → 訊息
Service → 服務
中文通常叫:
簡訊服務 / 手機簡訊
例如:
Your verification code is 123456.
送到 User 的 Phone Number。
SMS 常用於:
OTP
Security Alert
Delivery Update
Urgent Notification
剛剛出現一個縮寫:
OTP = One-Time Password
拆開:
One-Time → 一次性的
Password → 密碼
中文可以理解成:
一次性密碼 / 一次性驗證碼
例如:
Your verification code is 583921.
這個 Code 通常:
只能使用一次
或
短時間後失效
常用在:
Login Verification
Payment Verification
2FA
2FA = Two-Factor Authentication
拆開:
Two-Factor → 兩種驗證因素
Authentication → 身分驗證
中文:
雙因素驗證
例如 Login:
Factor 1:
Password
Factor 2:
SMS OTP
即使 Password 被偷,Attacker 還需要第二個 Factor。
Push Notification(推播通知):
System 把 Notification 推送到 User 的手機或裝置。
例如 iPhone 跳出:
Uber Eats
Your order is arriving.
User 不需要一直打開 App 問:
有沒有新通知?
而是 System 主動 Push。
In-App Notification(應用程式內通知):
User 打開 App / Website 後,在 App 裡看到的 Notification。
例如:
Facebook Notification Bell
GitHub Notifications
LinkedIn Notifications
可能顯示:
Alvin liked your post.
這和 Push Notification 不完全一樣。
Push:
可能在 App 沒打開時就出現在手機通知中心
In-App:
通常是 User 打開 App 後看到
Functional Requirement(功能需求):
System 必須提供哪些功能。
今天假設:
1. Send Email
2. Send SMS
3. Send Push Notification
4. Support User Notification Preferences
5. Support Scheduled Notifications
6. Support Retry
7. Track Delivery Status
Non-Functional Requirement(非功能需求):
System 要做到多快、多穩、多大、多可靠。
今天假設:
High Availability
Scalable
Reliable Delivery
Low Latency for urgent notifications
No unnecessary duplicate notifications
Provider failures should not crash the whole system
Reliable(可靠的):
System 在正常 Failure 情況下,仍盡量讓 Notification
最後能正確送出。
注意:
Reliable
不代表:
100% 永遠成功
因為可能:
Phone number invalid
Email address invalid
User blocks notifications
External provider unavailable
Device offline
比較合理的意思是:
System 應該能處理暫時性
Failure、Retry、記錄失敗狀態,而不是一失敗就直接把 Notification
丟掉。
假設:
100M Notifications/day
一天快速估成:
≈ 100K seconds
Average:
100M / 100K
≈ 1,000 Notifications/sec
假設 Peak:
10 × Average
Peak:
≈ 10K Notifications/sec
Notification Traffic 很可能不是平均分布。
例如:
9:00 AM
Marketing Campaign starts
突然:
10M Notifications
需要發送。
或者:
Concert tickets released
↓
Millions of users receive notification
所以:
Average Traffic
不能代表:
Peak Traffic
假設 Order Service 在 Order 建立後直接:
Order Service
↓
Send Email
程式概念:
createOrder()
sendEmail()
return success
看起來很簡單。
但是有問題。
Synchronous(同步):
前面的工作必須等待後面的工作完成,才能繼續。
例如:
Create Order
↓
Send Email
↓
Wait...
↓
Email sent
↓
Return Checkout Success
假設:
Create Order = 100 ms
Send Email = 2 seconds
User Checkout 可能要:
2.1 seconds
但 Email 並不是 Checkout 成功的必要條件。
Day 20 提過:
Critical Path(關鍵路徑):
User 的核心 Operation 成功前,必須完成的必要步驟。
Checkout:
Validate Order
Process Payment
Create Order
可能是 Critical Path。
但:
Send confirmation email
通常不是。
所以不應該讓:
Email Provider 很慢
直接讓:
Checkout 很慢
Asynchronous(非同步):
Caller 不需要等待所有後續工作完成,就可以先繼續。
例如:
Checkout
↓
Create Order
↓
Publish Event
↓
Return Success
Notification 可以稍後處理:
Event
↓
Notification System
↓
Send Email
這就是:
Asynchronous Processing(非同步處理)
這時可以使用:
Message Queue(訊息佇列)
拆開:
Message → 訊息
Queue → 佇列 / 排隊隊伍
可以想像成:
銀行抽號碼牌
Customer 不需要同時被處理。
先:
Take Number
↓
Wait in Queue
↓
Staff processes one by one
Message Queue 也是類似概念:
Producer
↓
Message Queue
↓
Consumer
Producer(生產者 / 訊息產生者):
產生 Message 並送進 Queue 的 Component。
例如:
Order Service
產生:
ORDER_CREATED
Event。
所以:
Order Service
=
Producer
Consumer(消費者 / 訊息處理者):
從 Queue 取出 Message 並處理工作的 Component。
例如:
Notification Worker
收到:
ORDER_CREATED
然後:
Send Email
所以 Notification Worker 是 Consumer。
Worker(工作處理程序):
專門從 Queue 或 Job System 取得工作並執行的 Process / Service。
例如:
Email Worker
SMS Worker
Push Worker
它們不一定直接處理 User HTTP Request。
主要工作是:
Background Processing
Order Service
↓
ORDER_CREATED
↓
Message Queue
↓
Notification Service
↓
Email Provider
↓
User
這樣:
Order Service
不用等待 Email 真正送完。
這裡出現一個重要概念:
Decoupling(解耦)
先理解:
Coupling(耦合):
兩個 Components 彼此依賴的程度。
如果:
Order Service
必須直接知道:
Email Provider API
SMS Provider API
Push Provider API
那兩邊 Coupling 很高。
Decoupling:
降低 Components 之間直接依賴的程度。
例如:
Order Service
↓
Publish ORDER_CREATED
它不用知道:
誰會 Send Email
誰會 Send SMS
誰會 Send Push
Notification System 自己處理。
Notification System 收到的資料可能像:
{
"userId": "123",
"type": "ORDER_SHIPPED",
"orderId": "A001"
}
注意:
最好不要讓所有上游 Service 都直接自己組完整 Email HTML。
Notification System 可以根據:
type
選擇 Template。
Template(範本 / 模板):
預先定義好的訊息格式,其中部分內容可以動態替換。
例如:
Hi {{name}},
Your order {{orderId}} has been shipped.
User:
name = Alvin
orderId = A001
最後:
Hi Alvin,
Your order A001 has been shipped.
Template 裡:
{{name}}
{{orderId}}
叫:
Placeholder(佔位符)
先保留一個位置,之後再放入真正的 Value。
例如:
Hello {{name}}
最後替換成:
Hello Alvin
可以有:
Template Service
負責:
Find Template
↓
Insert Variables
↓
Generate Final Message
例如:
ORDER_SHIPPED_EMAIL
對應某個 Email Template。
不是所有 User 都想收到所有 Notification。
例如 Alvin:
Order Update
Email = ON
SMS = OFF
Push = ON
這叫:
Notification Preference(通知偏好設定)
User 選擇哪些 Notification 類型可以透過哪些 Channels 收到。
Opt-in(選擇加入 / 同意接收):
User 主動選擇願意接收某類 Notification。
例如:
☑ Receive marketing emails
User 勾選後:
Marketing Email = ON
Opt-out(選擇退出 / 拒絕接收):
User 選擇不再接收某類 Notification。
例如:
Unsubscribe from marketing emails
這就是 Opt-out。
Transactional Notification(交易型 / 服務型通知):
和 User 的某個操作或帳戶事件直接相關的通知。
例如:
Order Confirmation
Password Reset
Security Alert
Payment Receipt
Marketing Notification(行銷通知):
主要用於 Promotion、Advertising、Campaign 的通知。
例如:
20% OFF today!
New products available!
這兩類 Notification 的 Priority、Preference、法律要求可能不同。
Notification Service 收到:
ORDER_SHIPPED
可以查 User Preference:
Email = ON
SMS = OFF
Push = ON
然後:
Email Worker → Send
SMS Worker → Skip
Push Worker → Send
這就是:
Channel Selection(通知管道選擇)
Notification System 通常不會自己直接連到全球所有 Email
Server、手機電信商或手機 OS。
通常會透過:
Provider(服務提供者)
提供實際外部傳送能力的第三方或平台。
例如概念上:
Notification System
↓
Email Provider
↓
Email Network
↓
User
或:
Notification System
↓
SMS Provider
↓
Telecom Network
↓
User Phone
手機 Push Notification 通常需要透過 Mobile Platform 提供的 Push
Service。
兩個很常見的縮寫:
APNs
FCM
這些一定要拆開。
APNs = Apple Push Notification service
拆開:
Apple
Push
Notification
service
中文可以理解成:
Apple 推播通知服務
它是 Apple 提供的 Push Notification Infrastructure。
簡化流程:
Your Backend
↓
APNs
↓
iPhone / iPad
也就是:
你的 Server 通常不是直接對 User 的 iPhone 發 Push,而是把 Push Request
交給 Apple 的 APNs,再由 Apple 傳到裝置。
FCM = Firebase Cloud Messaging
拆開:
Firebase → Google 的 App Development Platform 名稱
Cloud → 雲端
Messaging → 訊息傳遞
中文可以理解成:
Firebase 雲端訊息服務
它常用於 Android / App Push Messaging。
簡化:
Your Backend
↓
FCM
↓
User Device
APNs / FCM 要知道:
到底要送到哪一台 Device?
通常 App 會取得:
Device Token(裝置 Token)
Push Provider 用來識別某個 App Installation / Device Destination 的
Identifier。
可以簡化想像成:
Device Token
≈
Push Notification 的收件地址
例如:
User 123
↓
Device Token ABCXYZ...
Notification System 保存:
user_id → device_token
User 可能:
Uninstall App
Reinstall App
Change Device
Disable Notifications
因此 Device Token 可能:
Expired
Invalid
Changed
所以 Push Provider 回覆:
Invalid Token
時,System 應該更新或移除無效 Token。
我們可能需要保存:
Notification
├── notification_id
├── user_id
├── type
├── channel
├── status
├── created_at
├── scheduled_at
└── sent_at
另外:
User Preference
├── user_id
├── notification_type
├── email_enabled
├── sms_enabled
└── push_enabled
以及:
Device
├── user_id
├── device_token
├── platform
└── active
Status(狀態):
表示某個 Notification 現在處於哪個處理階段。
例如:
PENDING
QUEUED
SENT
FAILED
DELIVERED
這兩個不要混在一起。
SENT(已送出):
我們的 System / Provider 已經接受或送出 Notification。
DELIVERED(已送達):
Provider 確認 Notification 已經到達目標裝置或目的地。
例如:
Our System → Provider
成功:
SENT
不一定代表:
User Device 真正收到
所以:
Sent ≠ Delivered
而:
Delivered ≠ Read
User 收到,也不代表 User 已經打開。
Delivery Status(傳送狀態):
記錄 Notification 從建立到傳送過程目前進行到哪一步。
例如:
PENDING
↓
QUEUED
↓
SENT
↓
DELIVERED
Failure:
PENDING
↓
QUEUED
↓
FAILED
Provider 怎麼告訴我們:
Delivered
Failed
其中一種方式是:
Callback(回呼)
當某件事情完成後,對方主動呼叫我們提供的
Endpoint,把結果告訴我們。
例如:
Notification System
↓
SMS Provider
↓
Send SMS
稍後 Provider:
SMS Provider
↓
POST /delivery-status
↓
Notification System
告訴:
DELIVERED
這種 HTTP Callback 常叫:
Webhook
一個 System 在 Event 發生時,主動用 HTTP Request 通知另一個
System。
例如:
SMS delivered
↓
Provider
↓
POST our webhook
Webhook 和 Polling 不一樣。
Polling(輪詢):
我們自己固定一段時間去問對方:「有更新嗎?」
例如:
Every 10 seconds
↓
GET /status
Webhook:
有變化時
對方主動通知我們
Polling:
我們一直主動去問
外部 Provider 可能暫時失敗:
Timeout
503 Service Unavailable
Network Error
這些不一定代表永遠失敗。
所以可以:
Retry(重試)
Operation 失敗後,再嘗試一次。
例如:
Attempt 1 ❌
↓
Retry
↓
Attempt 2 ✅
Transient(暫時性的)
Transient Failure(暫時性失敗):
現在失敗,但稍後再試可能成功。
例如:
Temporary Network Error
Provider overloaded
Short timeout
這類 Failure 比較適合 Retry。
Permanent Failure(永久性失敗):
單純一直 Retry 通常不會解決的 Failure。
例如:
Invalid phone number
Invalid email address
Invalid device token
如果:
Phone Number 根本不存在
Retry 100 次也沒有意義。
所以 Retry 前要分辨 Failure Type。
如果 Provider Down:
Retry immediately
Retry immediately
Retry immediately
可能讓 Provider 更慘。
所以常使用:
Exponential Backoff(指數退避)
拆開:
Exponential → 指數成長
Backoff → 暫停 / 往後退一段時間再試
意思:
每次 Retry 失敗後,等待時間逐漸增加。
例如:
Retry #1 → wait 1 sec
Retry #2 → wait 2 sec
Retry #3 → wait 4 sec
Retry #4 → wait 8 sec
這不是唯一的時間設定,只是簡單例子。
假設 Provider 已經 Overloaded。
如果 10,000 Workers:
失敗
↓
立即 Retry
Provider 可能收到更多 Requests。
形成:
Provider overloaded
↓
Requests fail
↓
More retries
↓
Even more load
↓
Provider becomes worse
所以 Backoff 可以降低 Retry Pressure。
如果所有 Workers 都:
1 sec
2 sec
4 sec
8 sec
它們可能同一時間再次 Retry。
所以常加入:
Jitter(隨機抖動 / 隨機延遲)
在 Retry Wait Time 中加入一些 Randomness,避免大量 Clients
同時再次發 Request。
例如原本:
4 seconds
不同 Worker 可能:
3.7 sec
4.2 sec
4.8 sec
降低同時重試造成的 Traffic Spike。
假設:
Notification System
↓
SMS Provider
Provider 已經成功送出 SMS。
但是:
Response 在 Network 中丟失
我們看到:
Timeout
就以為:
Send failed
然後 Retry。
結果 User 收到:
Your OTP is 123456
Your OTP is 123456
兩次。
這叫:
Duplicate(重複)
這時需要一個非常重要的名詞:
Idempotency(冪等性)
中文:
冪等性
這個中文很不直覺。
最簡單理解:
同一個 Operation
執行一次或重複執行多次,最終效果應該盡量和執行一次相同。
例如:
Set notification status = SENT
做一次:
SENT
做五次:
SENT
最終狀態沒有變成五份。
Idempotency Key(冪等鍵):
用來識別「這是不是同一個 Operation」的 Unique Identifier。
例如:
notification_id = N123
第一次:
Send N123
Retry:
Send N123 again
System 可以知道:
這不是新的 Notification
這是同一個 Operation 的 Retry
Deduplication(去重):
拆開:
De- → 去除
Duplication → 重複
中文:
去除重複資料 / 重複工作
例如 Queue 裡意外出現:
N123
N123
N123
Consumer 可以檢查:
N123 already processed?
如果是:
Skip duplicate
兩者很接近,但可以先這樣理解:
Idempotency:
同一個 Operation 重做
→ 最終效果不要變成重複效果
Deduplication:
發現重複 Message / Request
→ 把重複的去掉或忽略
例如:
Message Queue delivers N123 twice
Deduplication:
第二個 N123 不處理
Idempotency:
就算第二個真的被處理
也不要造成第二次不應有的 Side Effect
Side Effect(副作用) 在 Software 裡不是「藥物副作用」。
它通常指:
Operation 對外部 State 造成的改變。
例如:
Send SMS
Charge Credit Card
Create Order
Update Database
Send Email
都是 Side Effects。
Retry Side Effect 時要特別小心。
Message Queue 常會談:
Delivery Semantics(傳遞語意)
簡單說:
Message System 對「Message 可能被送幾次」提供什麼保證。
常見:
At-most-once
At-least-once
Exactly-once
At-most-once:
拆開:
At most → 最多
Once → 一次
中文:
最多一次
意思:
Message 最多處理一次
可能:
0 次
或
1 次
但不會刻意重送。
Trade-off:
比較不容易 Duplicate
但可能丟 Message
At-least-once:
At least → 至少
Once → 一次
中文:
至少一次
意思:
System 會盡量確保 Message 至少被處理一次
因此可能:
1 次
2 次
甚至更多次
所以需要:
Idempotency
Deduplication
Exactly-once:
Exactly → 精確地
Once → 一次
中文:
恰好一次 / 精確一次
聽起來最完美:
每個 Message 永遠只產生一次效果
但在 Distributed System 中,要跨:
Queue
Consumer
Database
External Provider
Network
做到真正 End-to-End Exactly-once 非常困難。
因此實務上常見策略是:
At-least-once Delivery
+
Idempotent Processing
+
Deduplication
讓結果接近:
One logical effect
Message Queue 常出現:
ACK = Acknowledgement
中文:
確認 / 確認收到
Consumer 處理完 Message 後:
Consumer
↓
ACK
↓
Queue
意思:
這個 Message 我處理完成了。
Queue 才可以把它視為完成。
假設:
Consumer processes N123
↓
SMS sent successfully
↓
Consumer tries to ACK
↓
Network failure
Queue 沒收到 ACK。
Queue 可能認為:
N123 還沒完成
所以重新 Delivery。
這就是為什麼:
At-least-once
可能產生 Duplicate。
Retry 很多次還是失敗怎麼辦?
這時常使用:
DLQ = Dead Letter Queue
拆開:
Dead Letter → 無法正常送達 / 處理的訊息
Queue → 佇列
中文通常叫:
死信佇列
不要被名字嚇到。
可以理解成:
專門放「正常流程一直處理失敗」Message 的 Queue。
例如:
Notification
↓
Attempt 1 ❌
Attempt 2 ❌
Attempt 3 ❌
Attempt 4 ❌
↓
DLQ
如果一直 Retry:
Forever
可能浪費大量 Resources。
DLQ 可以:
隔離問題 Message
↓
不要一直阻塞正常 Queue
↓
讓 Engineer 之後 Investigate
例如:
Invalid Payload
Unexpected Provider Error
Bug
不是所有 Notification 都一樣重要。
Priority(優先順序):
決定哪些工作應該比較早被處理。
例如:
High Priority:
OTP
Security Alert
Payment Failure
Medium:
Order Update
Low:
Marketing Promotion
Priority Queue(優先佇列):
不是單純按照誰先進來,而是考慮 Priority 決定處理順序的 Queue。
一般 Queue:
First In
→ First Out
Priority Queue:
High Priority
→ process first
即使 Marketing Notification 比 OTP 更早進來,
OTP 也可能優先處理。
剛才提到:
FIFO = First In, First Out
拆開:
First In → 最先進來
First Out → 最先出去
中文:
先進先出
就像排隊:
Alvin
Bob
Charlie
Alvin 先到,所以 Alvin 先被服務。
有些 Notification 不是現在送。
例如:
Tomorrow 9:00 AM
→ Send meeting reminder
這叫:
Scheduling(排程)
指定某個工作在未來某個時間執行。
Database 可以保存:
scheduled_at
意思:
這個 Notification 預計什麼時間可以開始發送。
例如:
scheduled_at = 2026-10-05 09:00
Scheduler 到時間後:
Find due notifications
↓
Put into Queue
↓
Workers send
Due(到期 / 已到應執行時間)
例如:
scheduled_at <= now
表示:
這個 Notification 已經到可以發送的時間
如果 Taylor Swift 發一篇 Post:
100M Followers
要通知 Followers:
Taylor posted something new.
這時會遇到:
Fan-out
中文可以理解成:
把一個 Event 擴散成很多份工作。
例如:
1 Post Event
↓
100M Notification Jobs
這就是 Fan-out。
如果同步做:
Create Post
↓
Generate 100M Notifications
↓
Wait until complete
↓
Return success
User 可能等很久。
所以通常:
Create Post
↓
Publish Event
↓
Return
背景再慢慢:
Fan-out
↓
Generate Notification Jobs
↓
Queue
↓
Workers
Fan-out Worker:
把一個大型 Event 拆成大量 User-specific Notification Jobs 的
Worker。
例如:
CELEBRITY_POST_CREATED
↓
Fan-out Worker
↓
Follower 1 Notification
Follower 2 Notification
Follower 3 Notification
...
如果一次處理 100M Users,
可能會分成:
1000 users
1000 users
1000 users
...
這叫:
Batch(批次)
把大量 Data / Jobs 分成一批一批處理。
Batch Processing(批次處理):
一次處理一組 Data,而不是每筆都完全獨立立即處理。
例如:
Batch #1 → 1000 Users
Batch #2 → 1000 Users
Batch #3 → 1000 Users
可以幫助控制:
Memory
Database Load
Provider Load
假設 SMS Provider 只允許:
1,000 Requests/sec
但我們有:
10,000 SMS/sec
如果直接全部打過去:
Provider may reject requests
所以需要:
Rate Limiting(速率限制)
限制一段時間內最多執行多少 Requests。
例如:
SMS Provider
Limit = 1,000 RPS
Worker 就不能超過這個速度。
Throttling(節流):
主動降低處理速度,讓 Traffic 不超過 System 或 Provider
可以承受的範圍。
例如 Queue 裡:
100K Messages
不代表:
100K Messages 全部立刻送出去
可以:
Process 1K/sec
慢慢消化。
這裡有一個新的 Distributed System 名詞:
Backpressure(背壓 / 反壓)
中文名稱很抽象。
最簡單理解:
下游處理不過來時,System 要有方法讓上游不要無限制地繼續塞工作。
例如:
Producer = 10K msg/sec
Consumer = 1K msg/sec
如果一直這樣:
Queue grows
Queue grows
Queue grows
最後可能:
Memory / Storage exhausted
Backpressure 的核心思想:
Downstream is overloaded
↓
Slow down / limit upstream
Upstream(上游):
比較前面、產生 Request / Data 的 Component。
Downstream(下游):
接收並處理前面 Request / Data 的 Component。
例如:
Order Service
↓
Notification System
↓
SMS Provider
對 Notification System 而言:
Order Service = Upstream
SMS Provider = Downstream
假設:
SMS Provider A
完全 Down。
如果只有一個 Provider:
SMS Delivery stops
可以考慮:
Primary Provider
+
Secondary Provider
Failover(故障切換):
主要 Component 失敗時,把 Traffic / Work 切換到備用 Component。
例如:
SMS Provider A ❌
↓
Failover
↓
SMS Provider B ✅
Multi-Provider(多供應商):
同一種能力準備多個 Providers。
例如:
SMS
├── Provider A
└── Provider B
優點:
Higher Availability
Provider Failure Recovery
Trade-off:
More integration work
Different APIs
Different pricing
Different delivery behavior
More monitoring
不同 Provider API 可能完全不同。
例如:
Provider A:
sendSms(phone, message)
Provider B:
createMessage(to, body)
我們不希望 Business Logic 到處都是:
if provider == A
if provider == B
可以建立:
Adapter(轉接器 / 適配器)
把不同 External Interfaces 包裝成 System 內部一致的 Interface。
例如:
SMSProvider
|
├── ProviderAAdapter
└── ProviderBAdapter
上層只需要:
sendSMS()
不用知道 Provider 細節。
Routing(路由 / 分流):
決定一個 Request / Job 應該送到哪個 Destination。
例如:
US SMS
→ Provider A
Taiwan SMS
→ Provider B
或:
Provider A healthy
→ A
Provider A down
→ B
假設 Provider A 已經 Down。
如果每個 Request 還一直:
Call A
↓
Timeout 5 sec
↓
Retry
會浪費很多 Resources。
可以使用:
Circuit Breaker(斷路器)
這個名字來自電路中的 Breaker。
電路過載時:
Breaker opens
↓
Stop current
Software 裡:
當 Downstream 連續失敗太多次時,暫時停止呼叫它。
常見三個 States:
Closed
Open
Half-Open
正常狀態:
Requests can pass
Failure 太多:
Stop calling Provider
等待一段時間後:
Allow a few test requests
如果成功:
Back to Closed
如果還是失敗:
Back to Open
Timeout(逾時):
等待某個 Operation 超過指定時間後,就停止繼續等待。
例如:
SMS Provider
5 秒沒有 Response:
Timeout
沒有 Timeout 的話,
Worker 可能永遠卡住。
另一個常見縮寫:
TTL = Time To Live
拆開:
Time → 時間
To Live → 可以存活
中文:
存活時間 / 有效時間
意思:
一筆 Data / Message 最多有效多久。
例如 OTP Notification:
Your login code is 123456
如果:
OTP 5 minutes 後失效
那 Notification 30 分鐘後才送到已經沒有意義。
所以可以設定:
TTL = 5 minutes
如果 Queue 裡待太久:
now > created_at + TTL
就可以:
Drop / Expire
而不是繼續發送過期 Notification。
Expired(已過期):
超過有效時間,不應再按照原本用途處理。
例如:
Flash Sale ends at 12:00
12:30 才送:
Flash Sale starts now!
反而會傷害 User Experience。
Ordering(順序):
Messages 被處理的先後次序。
例如:
1. Order shipped
2. Order delivered
如果 User 收到:
Order delivered
↓
Order shipped
就很奇怪。
所以某些 Notification 需要考慮 Ordering。
Global Ordering(全域順序):
所有 Messages 都維持一個完整順序。
通常很昂貴,也不一定需要。
Per-User Ordering(每個 User 的順序):
只要求同一個 User 的相關 Messages 保持合理順序。
例如:
User A:
shipped → delivered
User B:
shipped → delivered
A 和 B 彼此誰先不重要。
這通常比 Global Ordering 更合理。
不要混淆。
Priority:
哪一類 Message 比較重要
Ordering:
同一組 Messages 的先後順序
例如:
Security Alert
Priority 很高。
而:
Order shipped → Order delivered
則有 Ordering Requirement。
現在把 Components 組起來:
Business Services
Order / Payment / Social / Security
↓
Events
↓
Message Queue
↓
Notification Service
/ | \
/ | \
↓ ↓ ↓
Email Queue SMS Queue Push Queue
↓ ↓ ↓
Email Worker SMS Worker Push Worker
↓ ↓ ↓
Provider Provider APNs / FCM
↓ ↓ ↓
User
因為:
Email
SMS
Push
可能有完全不同的:
Traffic
Provider
Rate Limit
Retry Policy
Priority
Cost
Failure Mode
例如:
SMS Provider Down
不應該讓:
Email Notification
也全部停止。
這是一種:
Failure Isolation(故障隔離)
Failure Isolation(故障隔離):
讓一個 Component 的 Failure 不要輕易擴散到其他 Components。
例如:
SMS Queue / Worker Failure
但:
Email Queue still works
Push Queue still works
這比全部共用同一條 Pipeline 更容易隔離問題。
Pipeline(處理管線 / 流程管線):
Data / Job 按照一連串 Steps 被處理的流程。
例如:
Event
↓
Preference Check
↓
Template
↓
Queue
↓
Worker
↓
Provider
↓
Delivery Status
這整條可以稱為 Notification Pipeline。
假設:
ORDER_SHIPPED
完整流程:
1. Order Service publishes ORDER_SHIPPED
2. Notification Service receives Event
3. Load User Preference
4. Determine Channels
5. Load Template
6. Generate Notification Jobs
7. Put Jobs into Channel Queues
8. Worker consumes Job
9. Check TTL
10. Call Provider
11. Provider returns result
12. Update Status
13. Retry temporary failures
14. Move repeated failures to DLQ
15. Receive Delivery Callback / Webhook
Notification System 也有 Security 問題。
例如 Notification 可能包含:
OTP
Password Reset Link
Bank Transaction
Personal Information
不能隨便被別人看到。
一個常見縮寫:
PII = Personally Identifiable Information
拆開:
Personally Identifiable
→ 可以識別某個人的
Information
→ 資訊
中文:
個人可識別資訊
例如:
Full Name
Email
Phone Number
Address
某些情況下:
Device identifiers
也可能屬於敏感識別資訊。
Sensitive Data(敏感資料):
如果被未授權的人看到、修改或洩漏,可能對 User 或 Organization
造成傷害的 Data。
例如:
OTP
Password Reset Token
Bank Information
Private Messages
所以 Logs 裡不要隨便寫:
Full OTP
Password
Access Token
Encryption(加密):
把可讀 Data 轉成 Ciphertext,持有正確 Key 才能還原。
Notification Data:
In Transit
At Rest
都可能需要考慮 Encryption。
In Transit(傳輸中):
Data 正在 Network 中移動。
例如:
Notification Service
↓
HTTPS
↓
Provider
HTTPS / TLS 可以保護傳輸中的 Data。
At Rest(靜態儲存中):
Data 正存在 Database、Disk、Storage 中,而不是正在 Network 傳輸。
例如:
Phone Number
Email Address
Notification History
可能需要適當的 At-Rest Protection。
Notification System 很複雜。
如果 User 說:
我沒有收到 OTP
你要能回答:
Event 有產生嗎?
Queue 有收到嗎?
Worker 有處理嗎?
Provider 有接受嗎?
Provider 有送達嗎?
Token 有效嗎?
Retry 了幾次?
是不是進 DLQ?
這就是 Observability 的價值。
可以監控:
Notifications Created/sec
Queue Backlog
Consumer Lag
Send Success Rate
Failure Rate
Retry Rate
DLQ Size
Provider Latency
Delivery Rate
p95 / p99 Latency
Backlog(積壓):
已經進入 Queue,但還沒有被處理完的工作數量。
例如:
Queue contains 1M messages
可能代表:
Producer 太快
Consumer 太慢
Provider 被 Rate Limited
Provider Failure
Consumer Lag(消費者延遲 / 消費落後量):
Consumer 處理進度落後 Producer 的程度。
簡單想:
Producer 已經產生到 Message #1,000,000
Consumer 只處理到 #800,000
Lag:
≈ 200,000 Messages
Lag 越來越大通常代表:
System processing cannot keep up
Delivery Rate(送達率):
成功送達的 Notifications 占應送 Notifications 的比例。
例如:
100,000 Notifications
95,000 Delivered
Delivery Rate:
95%
但不同 Channel 對 Delivered 的定義可能不同,要看 Provider 能提供什麼
Status。
End-to-End(端到端):
從整個流程最開始,到最終目標完成。
End-to-End Notification Latency:
Event Created
↓
Queue
↓
Notification Service
↓
Worker
↓
Provider
↓
Delivered
整段花的時間。
例如 Security OTP:
30 seconds
可能太慢。
所以不能只看:
Worker → Provider = 100 ms
還要看整條 Pipeline。
如果:
Queue Backlog ↑
而 Worker CPU / Provider Capacity 允許,
可以:
Add more Workers
例如:
Email Worker #1
Email Worker #2
Email Worker #3
...
這就是:
Horizontal Scaling(水平擴展)
增加更多 Machines / Processes 共同處理工作。
假設 Provider Limit:
1,000 RPS
你從:
10 Workers
增加到:
1,000 Workers
也不能無限制變快。
因為真正 Bottleneck 可能是:
Provider Rate Limit
所以:
Find Bottleneck
比:
Blindly add servers
重要。
Business Services
Order / Payment / Social / Security
↓
Events
↓
Message Queue
↓
Notification Service
/ | \
/ | \
↓ ↓ ↓
Preference Template Scheduler
Store Service |
\ | /
\ | /
↓ ↓ ↓
Notification Jobs
/ | \
↓ ↓ ↓
Email Queue SMS Queue Push Queue
↓ ↓ ↓
Email Workers SMS Workers Push Workers
↓ ↓ ↓
Email Provider SMS Provider APNs / FCM
↓ ↓ ↓
User
Failures:
Worker Retry
↓
Exponential Backoff + Jitter
↓
Repeated Failure
↓
DLQ
Delivery:
Provider
↓
Webhook / Callback
↓
Delivery Status Store
Observability:
Metrics + Logs + Traces
↓
Dashboard + Alerts
不是因為:
Notification System 標準答案就是這張圖
而是每個 Component 都在解決問題。
Problem:
Notification should not block business request
Solution:
Asynchronous Queue
Problem:
Email / SMS / Push have different failure and traffic patterns
Solution:
Separate queues and workers
Problem:
Users do not want every notification
Solution:
Store Opt-in / Opt-out preferences
Problem:
Message format should be managed consistently
Solution:
Reusable templates
Problem:
Temporary provider failures
Solution:
Retry
Problem:
Aggressive synchronized retries can overload provider
Solution:
Increase wait time + randomize retry timing
Problem:
Retry / queue redelivery can cause duplicate notification
Solution:
Identify same logical operation and suppress duplicate effects
Problem:
Some messages repeatedly fail
Solution:
Move them aside for investigation
Problem:
Provider has traffic limits
Solution:
Control send rate
Problem:
One provider can fail
Solution:
Secondary provider + failover
Problem:
User says notification never arrived
Solution:
Track every stage of pipeline
不要只說:
Because asynchronous is faster.
可以回答:
Notification delivery usually does not need to stay in the critical
path of the original business request. A Message Queue decouples the
business service from the notification workers, absorbs traffic
spikes, and allows workers to process notifications asynchronously. It
also gives us a place to implement retry and failure handling.
可以分層回答:
1. Timeout
2. Retry transient failures
3. Exponential Backoff
4. Add Jitter
5. Circuit Breaker
6. Keep jobs in Queue
7. DLQ after repeated failure
8. Failover to secondary provider if appropriate
9. Monitor provider failure rate
這比只回答:
Retry
完整很多。
可以說:
notification_id
↓
Idempotency Key
↓
Deduplication Store / Processing Record
然後解釋:
Because at-least-once delivery and network failures can cause
redelivery, consumers should be designed to be idempotent where
possible. We can use a unique notification ID to detect repeated
processing and avoid creating duplicate side effects.
因為不同 Channel 有不同:
Provider
Cost
Rate Limit
Latency
Retry Policy
Failure Mode
Priority
分開可以:
Scale independently
Fail independently
Monitor independently
這就是:
Failure Isolation
+
Independent Scaling
通常不合理。
例如:
OTP
User 正在 Login,可能希望:
幾秒內收到
Marketing Email:
晚 5 分鐘
通常影響很小。
所以可以:
High Priority
→ OTP / Security
Medium
→ Transactional Updates
Low
→ Marketing
但實際 Priority Policy 要根據 Product Requirements。
不要輕易說:
Yes.
更好的回答:
End-to-end exactly-once delivery is difficult in a distributed system,
especially when external providers are involved. A practical design
often uses at-least-once message delivery together with idempotent
processing and deduplication to prevent duplicate logical effects.
因為:
SENT
≠
DELIVERED
≠
READ
可能:
Provider accepted request
but device offline
Email entered spam
Push disabled
Device token invalid
SMS carrier delay
所以要明確定義每個 Status。
Requirements
□ Email?
□ SMS?
□ Push?
□ In-App?
□ Scheduled?
□ Analytics?
□ Delivery tracking?
Traffic
□ Notifications/day?
□ Average RPS?
□ Peak RPS?
□ Campaign spikes?
Channels
□ Different providers?
□ Different queues?
□ Different retry policy?
Preferences
□ Opt-in?
□ Opt-out?
□ Transactional?
□ Marketing?
Push
□ APNs?
□ FCM?
□ Device Token?
□ Invalid Token handling?
Reliability
□ Timeout?
□ Retry?
□ Transient vs Permanent Failure?
□ Exponential Backoff?
□ Jitter?
□ Circuit Breaker?
□ Multi-provider?
□ Failover?
Duplicates
□ At-least-once?
□ Idempotency?
□ Idempotency Key?
□ Deduplication?
Queue
□ ACK?
□ DLQ?
□ Backlog?
□ Consumer Lag?
□ Priority?
□ Ordering?
Scheduling
□ scheduled_at?
□ TTL?
□ Expired Notification?
Scale
□ Fan-out?
□ Batch?
□ Worker Scaling?
□ Provider Rate Limit?
□ Backpressure?
Security
□ PII?
□ Sensitive Data?
□ Encryption?
□ Authentication / Authorization?
Observability
□ Send Success Rate?
□ Delivery Rate?
□ Failure Rate?
□ Retry Rate?
□ Provider Latency?
□ End-to-End Latency?
□ DLQ Size?
1. Notification 是什麼?
2. Event 是什麼?
3. Channel 是什麼?
4. Email 是什麼?
5. SMS 全名與中文意思?
6. OTP 全名與中文意思?
7. 2FA 全名與中文意思?
8. Push Notification?
9. In-App Notification?
10. Reliable Delivery?
11. Synchronous?
12. Asynchronous?
13. Critical Path?
14. Message Queue?
15. Producer?
16. Consumer?
17. Worker?
18. Coupling?
19. Decoupling?
20. Template?
21. Placeholder?
22. Notification Preference?
23. Opt-in?
24. Opt-out?
25. Transactional Notification?
26. Marketing Notification?
27. Provider?
28. APNs 全名與用途?
29. FCM 全名與用途?
30. Device Token?
31. Status?
32. SENT vs DELIVERED?
33. Delivery Status?
34. Callback?
35. Webhook?
36. Polling?
37. Retry?
38. Transient Failure?
39. Permanent Failure?
40. Exponential Backoff?
41. Jitter?
42. Duplicate?
43. Idempotency 中文與意思?
44. Idempotency Key?
45. Deduplication?
46. Idempotency vs Deduplication?
47. Side Effect?
48. Delivery Semantics?
49. At-most-once?
50. At-least-once?
51. Exactly-once?
52. ACK 全名與中文意思?
53. ACK Failure 為什麼可能造成 Duplicate?
54. DLQ 全名與中文意思?
55. Priority?
56. Priority Queue?
57. FIFO 全名與中文意思?
58. Scheduling?
59. scheduled_at?
60. Fan-out?
61. Fan-out Worker?
62. Batch?
63. Batch Processing?
64. Rate Limiting?
65. Throttling?
66. Backpressure?
67. Upstream / Downstream?
68. Provider Failure?
69. Failover?
70. Multi-Provider?
71. Adapter?
72. Routing?
73. Circuit Breaker?
74. Closed / Open / Half-Open?
75. Timeout?
76. TTL 全名與中文意思?
77. Notification TTL?
78. Expired Notification?
79. Ordering?
80. Global Ordering?
81. Per-User Ordering?
82. Priority vs Ordering?
83. Failure Isolation?
84. Pipeline?
85. PII 全名與中文意思?
86. Sensitive Data?
87. Encryption?
88. Encryption in Transit?
89. Encryption at Rest?
90. Backlog?
91. Consumer Lag?
92. Delivery Rate?
93. End-to-End Latency?
94. Horizontal Scaling?
95. 為什麼 Notification System 適合 Message Queue?
96. Provider 掛掉怎麼辦?
97. 如何避免 Duplicate Notification?
98. 為什麼 Email / SMS / Push 分開處理?
99. OTP 和 Marketing Notification 為什麼 Priority 不同?
100. 為什麼 Exactly-once 很難?
Notification System 不是:
Business Service
↓
sendEmail()
而是要思考:
Business Event
↓
Message Queue
↓
Notification Service
↓
Preference
↓
Template
↓
Channel
↓
Worker
↓
Provider
↓
User
當 Provider 暫時失敗:
Retry
↓
Exponential Backoff
↓
Jitter
如果一直失敗:
DLQ
如果 Message 被重送:
Idempotency
+
Deduplication
如果大量 Users 同時要收到 Notification:
Fan-out
+
Batch
+
Queue
+
Horizontal