A bank runs a Foundry-based assistant for 4,000 staff. Usage is steady between 08:00 and 18:00, response latency must stay predictable at peak, and finance requires the monthly cost to be fixed in advance rather than varying with the number of tokens consumed. Which deployment option should you choose?
- A.
A fine-tuned copy of the model published on a standard deployment
- B.
A provisioned throughput deployment sized for the expected peak load
- C.
A batch deployment that processes queued requests offline and returns results later
- D.
A standard pay-as-you-go deployment that bills for each token consumed
Show answer
Answer: B
Provisioned throughput reserves capacity, which is what delivers predictable latency and a fixed monthly cost.
- A. Fine-tuning changes model behaviour, not capacity allocation or billing model, so neither constraint is met.
- B. Reserved capacity gives predictable throughput at peak and a committed price that does not vary with token volume.
- C. Batch returns results asynchronously and cannot meet interactive latency expectations for 4,000 staff.
- D. Per-token billing makes the monthly cost move with usage, which finance explicitly ruled out.
