iT邦幫忙

2026 iThome 鐵人賽

DAY 11
0

DynamoDB 的部分,要做 batch 的話,需要把 put_item 改用 batch_write_item ,才能把多筆資料一起塞到 table 裡。

def batch_write_to_dynamodb(items: List[Dict[str, Any]]) -> tuple:
              """
              Batch write items to DynamoDB with automatic retry for unprocessed items
              Returns: (success_count, failed_count)
              """
              if not items:
                  return 0, 0
              
              success_count = 0
              failed_count = 0
              
              # Split items into batches of 25 (DynamoDB limit)
              for i in range(0, len(items), BATCH_SIZE):
                  batch = items[i:i + BATCH_SIZE]
                  
                  try:
                      # Prepare batch write request
                      request_items = {
                          table.name: [
                              {'PutRequest': {'Item': item}} for item in batch
                          ]
                      }
                      
                      # Write batch with retry logic
                      response = dynamodb.batch_write_item(RequestItems=request_items)
                      success_count += len(batch)
                      
                      # Handle unprocessed items
                      unprocessed = response.get('UnprocessedItems', {})
                      retry_count = 0
                      max_retries = 3
                      
                      while unprocessed and retry_count < max_retries:
                          print(f"Retrying {len(unprocessed.get(table.name, []))} unprocessed items (attempt {retry_count + 1})")
                          response = dynamodb.batch_write_item(RequestItems=unprocessed)
                          unprocessed = response.get('UnprocessedItems', {})
                          retry_count += 1
                      
                      # Count final unprocessed items as failures
                      if unprocessed:
                          unprocessed_count = len(unprocessed.get(table.name, []))
                          failed_count += unprocessed_count
                          success_count -= unprocessed_count
                          print(f"⚠️ Failed to write {unprocessed_count} items after {max_retries} retries")
                  
                  except Exception as e:
                      failed_count += len(batch)
                      success_count -= len(batch)
                      print(f"❌ Error writing batch to DynamoDB: {str(e)}")
              
              return success_count, failed_count

在價錢上,DynamoDB 改用 batch 的方式並不會比較省錢,因為 DynamoDB 的計價是用 WRU ,一筆資料就算一個 WRU ,所以不管是不是用 batch ,花費的金額都是一樣的。

但除了省錢以外,實測10000個 Request ,可以從 Lambda 的 Duration 觀察到,右邊使用 batch 的方式處理資料, Duration 明顯比左邊好很多,這也是批次寫入的一個優點,不用都打很多次 API 增加處理時間。


上一篇
Day 10: 批次寫入資料 - S3
下一篇
Day 12: 用 AWS Glue Python Shell Job 把 S3 上的 JSON 檔轉成 Parquet 檔(上)
系列文
使用 Serverless 架構設計廣告點擊系統 共 16 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言