Every PPC manager runs tests. Few of them produce answers. The usual pattern is to change a bid, adjust a budget and add some negatives in the same week, see ACoS improve, and credit whichever change you were most excited about. The next time the same move fails, nobody knows why.

This playbook covers how to run a clean experiment in Amazon PPC: what makes a test clean, the two designs that work within the console's limits, how long to run them, and how to read the result without fooling yourself.

Why Amazon PPC is hard to test

In web advertising, a split test sends random visitors to version A or version B and compares them at the same time. Amazon Ads generally does not let you split the same search traffic between two versions of a campaign. Two campaigns targeting the same keyword compete with each other in the auction rather than dividing shoppers evenly.

Add attribution lag, weekly cycles, seasonality and competitors changing their own bids, and any single campaign's numbers move around a lot on their own. A test has to be designed to see through that noise. Amazon offers built-in experiments for some elements, such as listing content through Manage Your Experiments, and where one covers what you want to test, use it. For most bid, budget and schedule questions, you need one of the designs below.

The rules of a clean test

One change. Change one thing in the test group and nothing else. If you also need to add negatives, add them to both groups.

A control. Something comparable that does not change, so you can see what the market did on its own.

A decision in advance. Write down what you are testing, which metric decides it, how long it runs, and what you will do with each outcome. "If ACoS on the test group improves relative to control and total sales hold, roll it out to all generic campaigns."

A fixed length. Stopping when the numbers look good is the most common way to get a false result.

Design 1: test group and control group

Split comparable campaigns into two groups. Comparable means similar products, similar spend, similar targeting type. Apply the change to one group. Leave the other alone. Run both for the test period and compare the change in each group against its own baseline.

The comparison that matters is the difference between the two groups' changes. If the test group's ACoS fell by a few points and the control group's fell by the same amount, the change did nothing; the market moved. If the test group improved and the control did not, you have evidence.

This design works well for account-wide questions: dayparting, placement adjustments, bidding strategy, budget rules. It needs enough campaigns to split. With only two or three campaigns, the groups will differ too much to compare. Does dayparting work uses this design.

Design 2: alternating periods

When you cannot split campaigns, alternate the change on and off over time. One week on, one week off, repeated for at least four weeks each. Comparing several on weeks with several off weeks reduces the risk that one unusual week decides the result.

This design suits changes that take effect immediately and reverse cleanly, like a schedule or a placement adjustment. It does not suit changes with a slow effect, like a new keyword set that needs time to gather data. And attribution lag blurs the boundary between periods, so ignore the first few days of each.

What to measure

Pick one primary metric before you start. For most efficiency tests it is ACoS or ROAS. For growth tests it is ad sales or total sales at an acceptable ACoS. The metrics that matter covers which fits which question.

Then pick one or two guardrail metrics that must not get worse. A test that improves ACoS by cutting total sales is usually not a win. Check total sales, including organic, and impression share on your most important keywords.

Use enough data. A test decided on a handful of orders is a coin flip. If the campaigns in the test produce few orders per week, run longer or test on a larger group.

How long to run it

Four weeks is a good default. It covers each weekday several times, lets attribution settle for most of the period, and smooths out one-off events. Two weeks is the minimum for high-volume campaigns. Low-volume campaigns can need six to eight.

Avoid running tests across big seasonal shifts, such as the start of Q4 or a major sale event. The control group helps, but large events can affect the two groups differently.

Reading the result

Compare the change in the test group with the change in the control. If the difference is small relative to how much the numbers usually move from week to week, treat the result as inconclusive. That is a valid outcome. It means the change does not matter much, which is useful to know.

If the difference is clear and the guardrails held, roll the change out. Then keep watching for a few weeks, because tests on a subset do not always scale perfectly.

Record every test: hypothesis, design, dates, result, decision. A PPC change log is the right place. After a year, that record is one of the most valuable assets in the account, because it tells you what works for your products specifically.

Good first tests

Pausing or lowering bids in the weakest hours. Raising or lowering a top-of-search placement adjustment. Switching a set of campaigns from fixed bids to dynamic bids, down only. Moving converting search terms from auto into exact match. Each has a clear expected effect and a clear decision attached. The bid strategy guide covers the options for bidding tests.

Frequently asked questions

Can you A/B test Amazon PPC campaigns?

Not with a true randomized split in most cases, because Amazon does not let you divide the same traffic between two campaigns. You can run a before and after test with a control group, or split comparable products or campaigns into two groups. Amazon offers built-in experiments for some listing and ad elements where available.

How long should an Amazon PPC test run?

At least two full weeks, and usually four, so every weekday appears more than once and attribution has time to settle. Low-traffic campaigns need longer. Decide the length before starting and do not stop early because the first days look good or bad.

What should I test first in Amazon PPC?

Changes with a large expected effect and a clear decision attached, such as pausing weak hours, changing a placement adjustment, or switching bidding strategy on a set of campaigns. Small tweaks produce differences too small to detect reliably.


Off Hours lets you apply a rule to a test group of campaigns and leave the control untouched, with every change logged so the experiment stays clean. Start a free 14-day trial.