A cosmetics brand must choose between three Amazon Bedrock models to write product descriptions in its distinctive brand voice. Quality is subjective: marketing cares about tone, persuasiveness and whether claims stay within approved wording. The team has 600 product briefs, five experienced copywriters, and a written style rubric. Which TWO evaluation approaches should the team use to choose the model? (Select TWO.)
Choose 2.
- A.
Measure each model's perplexity on the brand's past catalogue text and select the model with the lowest perplexity, since it best predicts the brand's writing.
- B.
Score each model's outputs with an exact-match accuracy metric against one reference description per brief, because exact match measures brand voice objectively.
- C.
Compare the three models by their scores on a public benchmark for reasoning, and assume that the best reasoner also writes the most persuasive brand copy.
- D.
Run a model evaluation job that uses an LLM as a judge with a custom metric built from the written style rubric, scoring all three models on the 600 briefs.
- E.
Run a human-based evaluation job with the five copywriters as a private work team on a sample of briefs, rating each model's outputs against the same rubric.
Show answer
Answer: D, E
Subjective, rubric-driven quality is best measured with an LLM-as-a-judge custom metric at scale plus a human evaluation on a sample to confirm it.
- A. Perplexity measures how well a model predicts text, not the quality of the copy it generates, and it is not a Bedrock evaluation metric.
- B. Exact match penalises every valid rewording and cannot measure tone or persuasiveness.
- C. A reasoning benchmark does not measure brand voice or claim compliance.
- D. An LLM judge with a rubric-based custom metric scores subjective quality consistently across all 600 briefs.
- E. A human evaluation with the copywriters confirms that the automated scores match expert judgement.