ต้นทุน, Performance & Monitoringบทที่ 13/16 · ~9 นาที
Monitoring & Observability
ดูอะไร ที่ไหน — metric ของ Bedrock, invocation logs, tracing และ anomaly detection
| CloudWatch metric (AWS/Bedrock) | บอกอะไร |
|---|---|
| จำนวนการเรียก |
| latency ต่อการเรียก |
| ปริมาณ token → ใช้ประมาณต้นทุน, ตรวจ token burst |
| ถูก throttle (ต้อง backoff / cross-region / provisioned) |
| error 4xx / 5xx |
| จำนวนครั้งที่ guardrail เข้าแทรก |
เครื่องมือที่ต้องจับคู่ให้ถูก
- Model invocation logs → ดูเนื้อหา prompt/response แบบละเอียด วิเคราะห์ด้วย CloudWatch Logs Insights
- AWS X-Ray → trace request ข้าม API Gateway → Lambda → Bedrock หาจุดที่ช้า
- Agent trace → ดูขั้นตอน reasoning + tool calls ของ agent
- CloudWatch anomaly detection → แจ้งเตือนเมื่อ token usage หรือ latency ผิดปกติ
- Custom metrics (เช่น hallucination rate, user thumbs-down, cost ต่อ tenant) →
PutMetricData/ Embedded Metric Format - AWS Cost Anomaly Detection / Cost Explorer + cost allocation tags → ติดตามค่าใช้จ่ายต่อทีม
เช็กความเข้าใจ
ต้องการหาว่าส่วนไหนของ pipeline (API Gateway, Lambda, KB, Bedrock) ทำให้ latency สูง ควรใช้อะไร?
อ่านจบแล้ว? ลองทำข้อสอบเรื่องนี้เลย
การดึงความรู้ออกมาใช้ทันทีหลังอ่าน (retrieval practice) ช่วยให้จำได้นานขึ้นมาก