前面 17 天,我們已經學過 Load
Balancer、Database、Cache、Replication、Sharding、CDN、Message
Queue、Rate Limiter、Consistency、Reliability、Observability 與
Security。
但面試官如果說:
Design a URL Shortener.
不要立刻回答:
Redis!
Kafka!
Sharding!
Microservices!
因為你還不知道:
多少 User?
多少 Requests?
需要 Analytics 嗎?
短網址會過期嗎?
讀多還是寫多?
Latency 要多低?
System Design Interview 真正重要的是:
先理解 Problem,再根據 Requirement 做合理的 Engineering Decision。
今天使用這個流程:
Clarify Requirements
↓
Estimate Scale
↓
Define API
↓
Design Data Model
↓
Draw High-Level Design
↓
Find Bottlenecks
↓
Deep Dive
↓
Discuss Trade-offs
System Design Interview:
面試官給你一個大型 Software System Problem,請你說明會如何設計整個
System。
例如:
Design YouTube
Design Uber
Design URL Shortener
Design Chat System
Design Notification System
重點不是 45 分鐘內真的完成 Product,而是觀察:
怎麼理解 Requirement?
怎麼拆 Problem?
怎麼估 Scale?
怎麼選 Database?
怎麼處理 Traffic?
怎麼處理 Failure?
知道哪些 Trade-off?
System Design 通常沒有唯一答案。重點是把 Engineering Reasoning
說清楚。
Requirement:
System 必須滿足的需求。
例如 URL Shortener:
User 可以建立短網址
短網址可以 Redirect
System 要支援大量 URLs
Redirect 要很快
Requirement 不清楚,就不知道 System 到底需要做到什麼程度。
Functional Requirement:
System 必須提供哪些功能。
例如:
1. User 輸入 Long URL
2. System 產生 Short URL
3. User 開啟 Short URL
4. System Redirect 到 Long URL
簡單記:
Functional Requirement
→ System 要做什麼?
Non-Functional Requirement:
功能需要做到什麼品質或程度。
例如:
Redirect 要低 Latency
System 要 High Availability
支援大量 Traffic
Data 不應輕易遺失
簡單記:
Functional
→ 做什麼?
Non-Functional
→ 做得多快、多穩、多大、多安全?
一題 Design YouTube 可以包含:
Upload
Playback
Search
Recommendation
Comment
Like
Subscription
Live Streaming
Ads
不可能全部深入。
Scope:
這次 Design 包含哪些部分,以及哪些部分暫時不處理。
例如:
In Scope:
- Upload Video
- Watch Video
Out of Scope:
- Recommendation
- Comments
- Ads
In Scope = 這次處理。
Out of Scope = 這次先不處理。
Constraint:
設計時必須接受的限制或條件。
例如:
Must use existing PostgreSQL
Must support mobile clients
Data must stay in a region
Budget is limited
Response < 200 ms
它就像遊戲規則,Design 必須在這些限制內進行。
Assumption:
資訊不足時,暫時採用的合理假設。
例如:
Assume 100M Daily Active Users.
但不要偷偷假設。
可以說:
I'll assume around 100 million daily active users for estimation. Does
that sound reasonable?
這讓面試官可以確認或修正。
Clarify:
把模糊或不完整的 Requirement 問清楚。
例如:
Design a Chat System
先問:
1-to-1 chat?
Group chat?
Text only?
Images?
Read receipts?
Message history?
Maximum group size?
這叫 Requirement Clarification。
不同答案可能直接改變 Architecture。
Use Case:
User 使用 System 完成某個目標的情境。
URL Shortener:
Use Case 1:
Create Short URL
Use Case 2:
Open Short URL and Redirect
從 Use Case 開始,很容易找到 Functional Requirements。
DAU = Daily Active Users
一天內實際使用 Product 的不同 User 數量。
MAU = Monthly Active Users
一個月內實際使用 Product 的不同 User 數量。
例如:
Registered Users = 1B
MAU = 300M
DAU = 100M
Registered User 不代表每天都使用。
Scale:
System 要處理的使用量與資料量有多大。
可能包含:
Users
Requests
Traffic
Storage
Read Volume
Write Volume
100 Users 和 100M Users 的 Architecture 可能完全不同。
Capacity:
System 可以承受或需要提供多少處理能力。
Estimation:
估算。
所以:
Capacity Estimation = 粗略估算 Traffic、Storage、Bandwidth
等需求。
重點不是算到小數點,而是建立量級直覺。
字面是:
在信封背面快速算一下
意思:
不追求完全精準,快速估出大概量級。
例如:
100M Requests/day
一天:
86,400 seconds
平均:
100,000,000 / 86,400
≈ 1,157 RPS
≈ 1.2K RPS
Order of Magnitude:
數值大概位於哪個 10 的次方等級。
1K = 10^3
1M = 10^6
1B = 10^9
System Design 常更在乎:
1K RPS?
100K RPS?
10M RPS?
而不是 1,157 和 1,203 的差異。
RPS = Requests Per Second
每秒多少 Request。
QPS = Queries Per Second
每秒多少 Query / Operation。
簡化:
RPS → API / HTTP Request Rate
QPS → Query / Operation Rate
不同 Team 的使用方式可能略有不同。
假設:
86.4M Requests/day
平均:
86,400,000 / 86,400
= 1,000 RPS
這是 Average Traffic。
但 Real System 可能:
3 AM → 300 RPS
8 PM → 5,000 RPS
Peak Traffic:
Traffic 最高時段的負載。
只按照 Average Design,可能在 Peak 時壞掉。
如果沒有真實資料,Interview 可以假設:
Peak = 5 × Average
這個倍數叫:
Peak Factor
它不是固定答案。
Production 應根據 Historical Traffic;Interview 則要清楚說這是
Assumption。
URL Shortener:
Create Short URL
→ Write
Open Short URL
→ Read
如果:
Read : Write = 100 : 1
這就是 Read / Write Ratio。
它描述:
Read Traffic 與 Write Traffic 的比例。
Read-heavy:
Read Operations 遠多於 Write。
可能考慮:
Cache
Read Replica
CDN
Write-heavy:
Write Operations 很頻繁。
可能更需要關注:
Write Throughput
Partitioning
Sharding
Contention
Pattern:
反覆出現的形式或規律。
Traffic Pattern:
Traffic 隨時間或 Operation 呈現出的使用方式。
例如:
Read-heavy
Write-heavy
Morning Peak
Weekend Peak
Sudden Burst
Global Traffic
假設:
10M new records/day
500 bytes/record
每天:
10,000,000 × 500 bytes
≈ 5 GB/day
一年:
≈ 1.8 TB/year
Production 還可能包含:
Indexes
Replication
Backup
Metadata
Logs
所以實際 Storage 通常更多。
Metadata:
描述另一份 Data 的資料。
URL Record:
Long URL
Short Code
Created At
Expiration Time
Owner User ID
其中 Created At、Owner、Expiration 等就是描述 Record 的 Metadata。
照片的:
拍攝時間
相機型號
GPS
也是 Metadata。
Bandwidth:
一段時間內可以傳輸多少 Data。
假設:
1,000 RPS
Average Response = 100 KB
每秒:
1,000 × 100 KB
≈ 100 MB/sec
Image / Video System 特別需要注意 Bandwidth。
Latency Requirement:
Response Time 希望控制在哪個範圍。
例如:
p95 < 300 ms
Availability Requirement:
Service 需要多常保持可用。
例如:
99.9%
99.99%
Consistency Requirement:
Write 後,Read 需要多快看到最新資料。
Durability Requirement:
System 說 Data 已保存後,Failure 發生時是否仍應保留。
不同 Requirement 會影響不同 Architecture Decision。
API = Application Programming Interface
不同 Software Components 溝通時約定好的操作介面。
URL Shortener:
POST /urls
GET /{shortCode}
Frontend 不需要知道 Backend 內部怎麼實作,只需要知道 Request / Response
Contract。
API Design:
決定 System 提供哪些 Operations,以及 Request / Response 如何表示。
例如:
POST /urls
Request:
{
"longUrl": "https://example.com/very/long/path"
}
Response:
{
"shortUrl": "https://sho.rt/aB3x9"
}
Endpoint:
API 中一個特定可以被呼叫的入口。
例如:
POST /users
GET /users/123
DELETE /users/123
常見:
GET → 讀取 Resource
POST → 建立 Resource / 執行 Operation
PUT → 更新或取代
PATCH → 部分更新
DELETE → 刪除
實際 Semantics 仍取決於 API Design。
Resource:
System 裡可以被識別與操作的資料或概念。
例如:
User
Order
Product
Message
URL
REST = Representational State Transfer
初學者 Day 18 不需要背完整理論。
先記:
REST-style API 常用 HTTP Method + Resource URL 表達操作。
例如:
GET /users/123
POST /orders
DELETE /orders/456
Data Model:
System 有哪些主要 Data Entity、有哪些欄位,以及彼此有什麼關係。
URL Shortener:
URL
├── short_code
├── long_url
├── user_id
├── created_at
└── expires_at
Entity:
System 中需要保存與管理的一種主要資料物件。
例如:
User
Product
Order
Cart
Schema:
Data 在 Database 裡的結構定義。
Primary Key:
唯一識別每一筆 Record 的欄位或欄位組合。
Access Pattern:
Application 平常用什麼方式讀取或寫入 Data。
URL Shortener 最重要:
short_code
↓
long_url
Database Design 應該配合真正的 Access Pattern。
Query Pattern:
Database Query 平常怎麼查。
例如:
SELECT long_url
FROM urls
WHERE short_code = ?;
Access Pattern 更廣,可能包含:
Read by short code
Write new URL
Delete expired URL
List URLs by user
High-Level Design:
先畫主要 Components 與 Data Flow,不急著深入 Class 或 Function。
例如:
User
↓
Load Balancer
↓
URL Service
↓
Cache
↓
Database
第一版可以更簡單:
Client → Server → Database
再根據 Bottleneck 逐步演進。
Component:
System 中負責某類工作的主要部分。
例如:
Backend
Cache
Database
Message Queue
CDN
Data Flow:
Request / Data 在 Components 之間如何移動。
Architecture Diagram 不只是畫 Boxes,還要說清楚 Request 怎麼走。
Bottleneck:
限制整個 System Performance 或 Capacity 的主要地方。
例如:
10 Backend Servers
↓
1 overloaded Database
即使 Backend 很多,DB 仍可能限制整個 System。
Deep Dive:
從 High-Level Design 挑一個重要 Problem 深入分析。
URL Shortener 可能:
How to generate unique short codes?
How to handle huge redirect traffic?
時間有限,不需要每個 Component 都深入。
Hot Path:
最常執行、最重要或對 Performance 特別敏感的 Request Path。
URL Shortener 的 Redirect 可能是 Hot Path。
Critical Path:
Operation 成功前必須完成的步驟。
Analytics 不一定要放在 Create URL 的 Critical Path,可以非同步處理。
SPOF = Single Point of Failure
一個 Component 壞掉,就可能讓重要 Service 無法使用。
Failure Mode:
Component 可能用什麼方式失敗。
Database 不只會 Crash,也可能:
Slow Query
Connection Exhaustion
Disk Failure
Network Failure
Trade-off:
得到某個 Benefit 時,通常也需要付出某些 Cost。
例如 Cache:
Benefit:
Latency ↓
DB Load ↓
Cost:
Stale Data Risk ↑
Complexity ↑
不要只說:
Use Redis.
而要說:
為什麼需要?
解決什麼?
代價是什麼?
這是整個系列最重要的思考方式:
Problem
↓
Requirement
↓
Solution
↓
Trade-off
不要:
看到 System Design
→ 自動塞 Redis / Kafka / Sharding
Overengineering:
為實際不存在或還不重要的 Problem 加入過度複雜的 Design。
例如 100 Users 卻直接使用 100 Shards。
Underengineering:
Design 太簡單,無法滿足已知 Requirement。
例如:
1M RPS
99.99% Availability
卻只有:
1 Server
1 DB
No Backup
好的 Design 不是最複雜,而是符合 Requirement。
Iterative:
一輪一輪逐步改善。
例如:
Version 1
Client → Server → DB
↓ Server bottleneck
Version 2
Client → LB → Servers → DB
↓ DB read bottleneck
Version 3
Client → LB → Servers → Cache → DB
每個 Component 都有清楚加入原因。
題目:
Design a URL Shortener like TinyURL.
第一步不是 Redis。
而是:
Clarify Requirements
1. Create Short URL
2. Redirect Short URL
3. Optional Expiration
假設:
Out of Scope:
Detailed Analytics
Custom Alias
Low Redirect Latency
High Availability
Durable URL Mapping
Large Read Traffic
假設:
10M new URLs/day
1B redirects/day
Write:
10,000,000 / 86,400
≈ 116 RPS
≈ 120 RPS
Read:
1,000,000,000 / 86,400
≈ 11,574 RPS
≈ 12K RPS
所以:
Read : Write ≈ 100 : 1
如果:
Peak Factor = 5
Peak Read:
≈ 60K RPS
重點不是這些數字是標準答案,而是透過 Assumption 建立 Scale。
假設:
10M records/day
500 bytes/record
每天:
≈ 5 GB
一年:
≈ 1.8 TB raw data
還沒算:
Indexes
Replication
Backup
Create:
POST /urls
Request:
{
"longUrl": "https://example.com/article/123"
}
Response:
{
"shortCode": "aB3x9",
"shortUrl": "https://sho.rt/aB3x9"
}
Redirect:
GET /aB3x9
Redirect:
Server 告訴 Client:「Resource 在另一個 URL,請去那裡。」
GET sho.rt/aB3x9
↓
Server
↓
Location: example.com/article/123
↓
Browser visits Long URL
HTTP 常使用 3xx Status Code 表示 Redirect。
301 Moved Permanently:
比較偏向永久 Redirect。
302 Found:
比較偏向暫時 Redirect。
Browser、Cache、Search Engine 對不同 Redirect Status
可能有不同處理方式。
所以不要死背:
URL Shortener 一定用 301
要依 Analytics、Caching 等 Requirement 選擇。
URL
├── short_code
├── long_url
├── created_at
└── expires_at
最重要 Access Pattern:
short_code → long_url
所以 short_code Lookup 必須有效率。
User
↓
URL Service
↓
Database
Create:
POST /urls
↓
Generate Short Code
↓
Database
↓
Return Short URL
Redirect:
GET /aB3x9
↓
Database
↓
Long URL
↓
Redirect
我們估算:
12K average read RPS
60K peak read RPS
如果全部打 Database,DB 可能成為 Read Bottleneck。
而 URL Mapping:
short_code → long_url
非常適合 Cache。
因此:
User
↓
Load Balancer
↓
URL Service
↓
Redis
↓ Cache Miss
Database
這是 Cache-Aside Pattern。
不是因為:
Redis 很熱門
而是:
Read : Write ≈ 100 : 1
Popular URLs repeatedly requested
所以:
Cache
→ Reduce DB Load
→ Reduce Latency
Trade-off:
Cache Invalidation
Stale Data
Memory Cost
Cache Failure
一台 Service 不夠:
User
↓
Load Balancer
↓
URL Service #1
URL Service #2
URL Service #3
Backend 可以盡量 Stateless,方便 Horizontal Scaling。
如果 Data 持續增加:
Millions → Billions
先看真正 Bottleneck。
Read Problem:
Cache
Read Replica
Storage / Write Problem:
Partitioning
Sharding
不要看到 Billion 就自動回答 Sharding。
如果:
URL A → abc123
URL B → abc123
System 不知道 abc123 應該去哪裡。
Collision:
兩個不同 Input 最後得到相同 Identifier / Result。
Uniqueness:
某個 Value 不和其他 Record 重複。
Short Code 通常需要維持 Uniqueness。
Random Code 可能:
Generate Code
↓
Check
↓
Exists?
├─ Yes → Generate Again
└─ No → Save
Collision Handling:
發生重複時,System 如何處理。
Hot Key:
某個 Cache Key 被非常大量 Requests 同時存取。
例如 Celebrity 分享:
sho.rt/viral
突然得到極大量 Traffic。
即使有 Cache,單一 Cache Node 也可能受到很大壓力。
Happy Path:
一切正常時最主要、最順利的流程。
例如:
Valid Short URL
→ Cache Hit
→ Redirect Success
Edge Case:
特殊或邊界情況,需要正確處理。
例如:
Collision
Expired URL
Invalid URL
Deleted URL
Extremely Popular URL
Error Handling:
System 遇到 Error / Abnormal Condition 時如何處理。
例如:
Unknown Short Code
→ 404 Not Found
Database Timeout:
Retry?
Fallback?
Return Error?
需要根據 Requirement 決定。
User
↓
DNS / CDN
↓
Load Balancer
/ \
↓ ↓
URL Service URL Service
\ /
\ /
Redis
↓
Database
Redirect:
GET /aB3x9
↓
URL Service
↓
Redis
├─ Hit → Long URL
└─ Miss → Database → Cache
↓
HTTP Redirect
System Design Interview 不只考 Architecture Knowledge。
Communication:
讓面試官知道你正在想什麼,以及為什麼做這個 Decision。
例如:
Since redirect traffic is much higher than URL creation traffic, I'll
optimize the read path first.
這比安靜畫 15 分鐘更清楚。
Signposting:
說明複雜內容時,先告訴對方現在在哪一步、下一步要做什麼。
例如:
First, I'll clarify the requirements. Then I'll estimate traffic and
storage. After that, I'll propose a high-level design.
面試官更容易跟上你的思考。
Design Decision:
多個方案中,根據 Requirement 選擇一個方案。
例如:
Decision:
Use Cache
Reason:
Read-heavy
Trade-off:
Stale Data / Complexity
Alternative:
可以解決相同 Problem 的另一個方案。
例如:
SQL
vs
Key-Value Store
不一定要把所有 Alternatives 都講完,但重要 Decision 可以說明為什麼選 A
而不是 B。
Premature:
過早的。
Premature Optimization:
還沒有證據顯示某處是重要 Bottleneck,就過早投入很多 Complexity 優化。
例如還不知道 Traffic 就:
Shard into 100 databases
更好的流程:
Requirement
↓
Estimate
↓
Find likely bottleneck
↓
Optimize
這裡的 Framework 不是 React / Spring。
它指:
一套幫助你有系統思考 Problem 的結構。
例如:
Requirements
↓
Scale
↓
API
↓
Data
↓
Architecture
↓
Bottleneck
↓
Trade-off
不是每題一定完全照順序,而是在不知道下一步時幫助整理思路。
1. Clarify Requirements
2. Define Scope
3. Functional Requirements
4. Non-Functional Requirements
5. Estimate Scale
6. Estimate Traffic / Storage
7. Read / Write Pattern
8. Define APIs
9. Define Data Model
10. High-Level Design
11. Walk Through Data Flow
12. Find Bottlenecks
13. Deep Dive
14. Discuss Failure
15. Discuss Security / Reliability
16. Discuss Trade-offs
17. Summarize
Kafka!
Redis!
MongoDB!
卻回答不出為什麼。
這可以叫 Technology Dump:
列出很多 Technology,卻沒有說明它們解決什麼 Problem。
100 Users
和:
1 Billion Users
Architecture 可能完全不同。
還要問:
Cache fails?
DB slow?
Backend dies?
Traffic spikes?
不要只說:
Redis is fast.
要說 Benefit 和 Cost。
Components 越多不代表 Design 越好。
不要把自己猜的數字當成 Fact。
不要直接:
Add Redis.
先問:
Which query?
Read or Write?
Traffic increased?
Missing Index?
Connection Pool full?
Lock contention?
再依 Evidence 選:
Index
Query Optimization
Cache
Read Replica
Partition
Sharding
先問:
Add Product?
Remove Product?
Change Quantity?
Guest Cart?
Checkout included?
再問:
How many users?
How long should cart persist?
Can cart be temporarily stale?
Data Model:
User
Cart
CartItem
Product
Access Pattern:
Get cart by user_id
Add item
Remove item
Update quantity
然後才討論 SQL、NoSQL、Redis。
Because [Problem / Requirement],
I would consider [Solution],
which helps [Benefit],
but the trade-off is [Cost / Risk].
例如:
Because redirect traffic is read-heavy and popular mappings are
repeatedly requested, I would consider Redis as a cache. This reduces
database load and latency, but introduces cache invalidation and
stale-data concerns.
這是在展示:
Engineering Reasoning
而不是背 Architecture。
Requirements
□ Functional?
□ Non-Functional?
□ Scope?
□ Constraints?
□ Assumptions?
Scale
□ DAU / MAU?
□ Average RPS?
□ Peak RPS?
□ Read / Write Ratio?
□ Storage?
□ Bandwidth?
Quality
□ Latency?
□ Availability?
□ Consistency?
□ Durability?
□ Security?
Interface
□ APIs?
□ Resources?
Data
□ Entities?
□ Schema?
□ Access Patterns?
Architecture
□ High-Level Design?
□ Data Flow?
□ Hot Path?
□ Critical Path?
Reliability
□ SPOF?
□ Failure Modes?
□ Monitoring?
Discussion
□ Bottleneck?
□ Deep Dive?
□ Alternatives?
□ Trade-offs?
1. Requirement?
2. Functional Requirement?
3. Non-Functional Requirement?
4. Scope?
5. Constraint?
6. Assumption?
7. Requirement Clarification?
8. Use Case?
9. DAU / MAU?
10. Scale?
11. Capacity Estimation?
12. Back-of-the-Envelope Estimation?
13. Order of Magnitude?
14. RPS / QPS?
15. Average vs Peak Traffic?
16. Peak Factor?
17. Read / Write Ratio?
18. Read-heavy / Write-heavy?
19. Traffic Pattern?
20. Storage Estimation?
21. Metadata?
22. Bandwidth?
23. Latency Requirement?
24. Availability Requirement?
25. Consistency Requirement?
26. Durability Requirement?
27. API / Endpoint?
28. HTTP Method?
29. Resource / REST?
30. Data Model?
31. Entity / Schema / Primary Key?
32. Access Pattern / Query Pattern?
33. High-Level Design?
34. Component / Data Flow?
35. Bottleneck?
36. Deep Dive?
37. Hot Path / Critical Path?
38. SPOF / Failure Mode?
39. Trade-off?
40. Overengineering / Underengineering?
41. Iterative Design?
42. Redirect / 301 / 302?
43. Collision / Uniqueness?
44. Collision Handling?
45. Hot Key?
46. Edge Case / Happy Path?
47. Error Handling?
48. Communication / Signposting?
49. Design Decision / Alternative?
50. Premature Optimization?
51. Framework?
52. Technology Dump?
53. 為什麼不能一開始就選 Redis / Kafka?
54. Database 很慢怎麼分析?
55. 如何回答「為什麼用 Redis?」
System Design Interview 最重要的不是:
知道最多 Technology
而是:
根據 Requirement 做合理的 Engineering Decision。
實用流程:
Clarify Requirements
↓
Define Scope
↓
Estimate Scale
↓
Define API
↓
Design Data Model
↓
High-Level Design
↓
Walk Through Data Flow
↓
Find Bottleneck
↓
Deep Dive
↓
Discuss Failure
↓
Discuss Trade-offs
核心:
Problem
↓
Requirement
↓
Solution
↓
Trade-off
好的 System Design,不是 Diagram 裡有多少
Boxes,而是你能不能清楚解釋每一個 Box 為什麼存在。
今天已經開始使用:
DAU
RPS
Peak Traffic
Read / Write Ratio
Storage
Bandwidth
下一篇:
Day 19|Back-of-the-Envelope Estimation:System Design 面試中的
Traffic、Storage、Bandwidth 到底怎麼估?
會從最基礎解釋:
Byte
KB
MB
GB
TB
Second / Day Conversion
RPS
QPS
Peak Traffic
Storage Growth
Bandwidth
Memory Estimation
Cache Size
Replication Factor
並用 URL Shortener、Chat System、Photo / Video System 一步一步練習。