Exam CCAR-P Topic 1 Question 30 Discussion
Actual exam question for Anthropic's CCAR-P exam
Question #: 30
Topic #: 1
Question #: 30
Topic #: 1
You are running a controlled experiment to compare two prompts and must complete the design steps before executing the experiment.
Which two steps must be completed BEFORE running the experiment with random assignment? (Select two.) Each correct answer presents part of the solution.
Which two steps must be completed BEFORE running the experiment with random assignment? (Select two.) Each correct answer presents part of the solution.
Suggested Answer: A,C Vote an answer
A controlled prompt experiment must begin with a falsifiable hypothesis and a predefined primary metric.
Option C prevents the team from examining results first and then selecting whichever metric makes the candidate look successful. The metric might measure task accuracy, rubric score, citation validity, escalation rate, latency, cost, or another criterion directly connected to the hypothesis.
Option A determines whether the experiment can detect a practically meaningful improvement. The minimum detectable effect expresses the smallest difference worth acting upon, while the power calculation determines the required sample size. Without this step, the experiment may be too small to detect a real improvement or unnecessarily large and expensive.
Random assignment should then distribute representative traffic between the control and candidate prompts while controlling model version, retrieval configuration, tool availability, and other confounding variables.
Options B and D occur after data collection. Option E follows the completed analysis and decision. The team should also define significance thresholds, stopping rules, guardrail metrics, exclusion criteria, and treatment of repeated observations before launch.
Study Guide references/topics: Prompt A/B testing; hypothesis definition; primary metrics; minimum detectable effect; statistical power; random assignment; decision sequencing.
Option C prevents the team from examining results first and then selecting whichever metric makes the candidate look successful. The metric might measure task accuracy, rubric score, citation validity, escalation rate, latency, cost, or another criterion directly connected to the hypothesis.
Option A determines whether the experiment can detect a practically meaningful improvement. The minimum detectable effect expresses the smallest difference worth acting upon, while the power calculation determines the required sample size. Without this step, the experiment may be too small to detect a real improvement or unnecessarily large and expensive.
Random assignment should then distribute representative traffic between the control and candidate prompts while controlling model version, retrieval configuration, tool availability, and other confounding variables.
Options B and D occur after data collection. Option E follows the completed analysis and decision. The team should also define significance thresholds, stopping rules, guardrail metrics, exclusion criteria, and treatment of repeated observations before launch.
Study Guide references/topics: Prompt A/B testing; hypothesis definition; primary metrics; minimum detectable effect; statistical power; random assignment; decision sequencing.
by Alma at Sep 12, 2026, 10:03 PM
0
0
0
10
Comments
Upvoting a comment with a selected answer will also increase the vote count towards that answer by one. So if you see a comment that you already agree with, you can upvote it instead of posting a new comment.
Report Comment
Commenting
You can sign-up / login (it's free).