Score the test backlog — the impact, the effort, the confidence — so the team runs the experiments that move the metric, not the ones they feel like running.
Score our 12-experiment backlog and rank the top 3.
The backlog scored and ranked in an hour. The top 3 shipped in order — 2 of 3 moved the metric, which is the hit rate the score predicted. The team stopped running the loudest idea and started running the highest-leverage one, which is the point of the score.
An experiment prioritization template is the score that turns a test backlog into a ranked list. It fixes the shape — the impact, the effort, the confidence — so the team runs the experiments that move the metric, not the ones they feel like. A backlog without a score is a backlog that runs the loudest idea; a backlog with a score is a backlog that runs the highest-leverage one.
A reusable scoring shape for a test backlog — the impact, the effort, the confidence — that ranks the experiments. The template fixes the shape so the team runs the experiments that move the metric, not the ones they feel like.
By the metric the experiment moves — revenue, conversion, retention — not by the page the experiment is on. Impact is the part that earns the score, because an experiment that moves a vanity metric is an experiment that does not move the business.
The engineering and design hours the experiment takes to ship. Effort is the part that earns the rank, because an experiment that takes a month for a 2% lift is an experiment that loses to three experiments that take a week each for a 5% lift.
The odds the experiment will move the metric — based on prior tests, user research, or industry proof. Confidence is the part that earns the rank, because an experiment with high impact and low confidence is an experiment that wastes the team’s time.
The agent reads your test backlog and the metrics, scores each experiment on impact, effort, and confidence, and ships a ranked list. It pairs with the experiment hypothesis brief and the experiment log summary templates.
The experiments the team is considering.
Agent fills impact, effort, confidence per experiment.
By impact × confidence ÷ effort.
Ship the highest-ranked experiment first.
Scores on the metric, not the page
Effort counts the hours, not the guess
Confidence counts the odds, not the hope
Pairs with the hypothesis brief and log templates
Score each experiment on impact (the metric it moves), effort (the hours to ship), and confidence (the odds it moves the metric), then rank by impact × confidence ÷ effort. The team runs the top of the backlog first, which is the highest-leverage test, not the loudest idea.
Impact × confidence ÷ effort. The framework is simple enough that the team can score a backlog in an hour, and rigorous enough that the top of the list is the highest-leverage test. Frameworks that add more factors tend to add more opinions, not more rigor.
By the metric the experiment moves — revenue, conversion, retention — not by the page the experiment is on. Impact is the part that earns the score, because an experiment that moves a vanity metric is an experiment that does not move the business.
By the odds the experiment will move the metric — based on prior tests, user research, or industry proof. Confidence is the part that earns the rank, because an experiment with high impact and low confidence is an experiment that wastes the team’s time.