Cost Optimization
- Prefer
-mini/-nano/-flashmodels for simple tasks - Use a larger-context model for long documents, a cheaper one for short tasks
- Set a budget cap to prevent overspending
- Reuse prompt context and enable caching in batch workloads
-mini / -nano / -flash models for simple tasks