Cost & Performance Optimization
ลด token, cache อย่างฉลาด, เลือกโมเดลตามความยาก และจัดการ throughput
| เทคนิค | ลดอะไร | ใช้เมื่อ |
|---|---|---|
Prompt caching | ต้นทุน (สูงสุด ~90%) และ latency (สูงสุด ~85%) ของ prefix ที่ซ้ำ | system prompt/เอกสารยาวที่ใช้ซ้ำทุก request |
Semantic caching | การเรียก FM ซ้ำสำหรับคำถามที่ความหมายเหมือนกัน | FAQ / คำถามซ้ำๆ — เก็บ embedding+คำตอบใน ElastiCache/MemoryDB/OpenSearch |
Exact-match cache | request ที่เหมือนกันทุกตัวอักษร | hash ของ request → DynamoDB/ElastiCache |
Model cascading / tiering | ต้นทุนต่อคำถาม | ส่งคำถามง่ายไปโมเดลเล็ก ยากค่อยใช้โมเดลใหญ่ |
Intelligent Prompt Routing | ต้นทุน โดยให้ Bedrock เลือกโมเดลในตระกูลเดียวกันให้ | อยากได้ tiering แบบ managed |
Batch inference | ราคา ~50% | งาน offline |
Context pruning / summarize history | input tokens | แชตยาว context ล้น |
| output tokens | คำตอบยาวเกินจำเป็น |
Latency-optimized inference | latency | แอปที่ time-sensitive (บางโมเดล/Region) |
Provisioned Throughput | throttling, latency ไม่คงที่ | traffic สูงคงที่ ต้องการ capacity การันตี |
แชตบอตส่ง system prompt + คู่มือยาว 20,000 tokens ทุก request ทำให้ค่าใช้จ่ายสูง ควรเปิดใช้อะไรก่อน?
การดึงความรู้ออกมาใช้ทันทีหลังอ่าน (retrieval practice) ช่วยให้จำได้นานขึ้นมาก