Google Ads Experiments: Hypothesis, Duration and Analysis

Zwei vergleichbare Materialproben in einer historischen Architekturwerkstatt mit gemeinsamem Prüfrahmen und Beobachtungsjournal als Bild für sorgfältig geplante Google-Ads-Experimente

Google Ads · Experiment planning · Updated: 14 September 2026

1. Start an experiment with the decision that follows

Google Ads experiments let you compare a specific change with your current campaign settings at the same time. The comparison becomes meaningful when the hypothesis, measurement, assignment, duration and decision rule fit together before launch. A green result indicator does not, by itself, answer a business question.

You may want to try a different bidding strategy, simplify a form or test a new targeting approach. The first question is: what decision would a convincing result lead to? If you have to adopt the variant regardless of the outcome, a monitored, controlled rollout may be more appropriate. An experiment plan is worthwhile when the result can inform a real choice.

The result must support an action

This guide focuses on custom Search experiments for small and medium-sized businesses. Its practical tool is the Google Ads experiment protocol: a template connecting the research question, measurement agreement, calendar and approval. You can copy the tables into your own document. The suggested fields and business thresholds are an editorial planning aid, not an additional Google feature.

There are four clear outcomes: adopt the variant, reject the variant, leave the question unresolved, or declare the experiment uninterpretable because of a disruption. The last two are useful outcomes too, if they prevent an uncertain finding from becoming the permanent campaign standard.

2. Choose the experiment to match the question

A before-and-after comparison mixes your change with everything that happens between the two periods: demand, competition, public holidays or sales staffing. Running the comparison simultaneously reduces these differences over time. It does not resolve an unclear measurement definition or repair a technically broken variant.

Google offers several experiment options, including ad variations and different campaign experiments. The overview of experiment types helps you select an appropriate option. A statement about Search cannot be applied to Performance Max, Demand Gen or video without checking it. Their requirements, reporting and options for applying results may differ.

First check whether you can serve the change separately

The active original campaign must support the intended comparison and continue running throughout it. Shared budgets or unsupported legacy ad formats can prevent setup. Resolve these obstacles before approving the schedule. Delayed ad approval can also mean that the planned launch is not yet the actual start of a comparable test.

For a landing page change, you need two versions that can each be reached reliably. Check that your chosen experiment type assigns them correctly and that an ad edit does not accidentally affect both variants. A new form, new headline and different offer together would be a package test: you could examine their combined effect, but could not identify any one component as the cause.

3. Write a testable hypothesis

“We are testing optimisations” leaves both the change and the definition of success open. A useful hypothesis connects an observed obstacle with one planned change, a plausible mechanism and a commercially relevant metric. The reasoning explains your expectation; it does not yet prove it.

A concrete teaching example for a B2B supplier

The fictional business Rheinbogen Industrietechnik receives enquiries about industrial maintenance. It suspects that two additional mandatory fields discourage suitable prospects. The experiment will make only these two fields optional: four mandatory fields instead of six, with the same offer and an unchanged definition of a sales-qualified enquiry. This is an invented teaching example, not Salestudia client data.

Hypothesis: “Four mandatory fields instead of six will reduce cost per qualified enquiry by at least 10%, because fewer suitable prospects abandon the form. Qualified enquiry volume and the sales team's ability to process those enquiries must remain within our agreed guardrails.”

Management chooses the 10% threshold itself. With a mature baseline CPA of €200, it represents at least €20 less in advertising costs per qualified enquiry. Assuming the same volume of 100 enquiries, that would be €2,000. Constant volume is an assumption for the comparison, not a forecast. If implementation and additional checking consume that benefit, the threshold has been set too low.

4. Complete the Google Ads experiment protocol

Create an approved version before setting up the experiment. The second column contains entries to replace with details of your own test. The example column shows how the parts connect without prescribing universal limits. Then add the metric agreement, calendar and final decision described in the following sections.

One document connects the business, account and analysis

ComponentYour entryRheinbogen exampleEvidenceOwnerApproval check
Identity[ID, version, date, campaigns]RB-FORM-01, version 1, maintenance Search campaignAccount names and campaign IDsAccount managerExperiment route and original campaign clear?
Baseline finding[Observation, reference period, season]Form abandonment; mature CPA of €200Form analysis and mature Ads dataAnalyticsNo unresolved measurement disruption?
Question[Change, mechanism, decision]Two mandatory fields become optionalDocumented form variantsWebsite ownerOnly the planned change?
Comparison[Version A, version B, implementation]Six versus four mandatory fields; everything else unchangedSeparately tested page versionsWebsite ownerAre the two arms actually different?
Measurement[Action, formula, filters, attribution, time zone, currency]CPA for a qualified enquiry, in eurosCRM definition and Ads configurationSales and analyticsConsistent across both arms?
Assignment[Method, unit, split, budget]Cookie-based, planned 50/50 splitSaved setupAccount managerShared resources checked?
Economics[Threshold and rationale]At least 10% lower CPA€200 baseline CPA, estimated effortManagementBenefit remains after additional work?
Guardrail metrics[Limits, source, maturity, owner]Volume, ability to process enquiries and experiment costsCompleted metric agreementSales and account managerSpecific limits approved?
Feasibility[Volume per arm, variability, precision, time and cost limits]Review mature history before approving the calendarVolume forecast with stated assumptionsAnalyticsCan this data budget support the question?
Statistics[Method, level and, where relevant, power assumptions]Experiment report; 95% chosen in advanceReport setting and planning noteAnalyticsDoes the method suit the metric?
Calendar[Start, ramp-up, measurement phase, cut-off]Only after checking volume and maturityHistory and conversion delayAnalyticsMaximum affordable duration included?
Intervention rules[Maintenance, log, stopping, extension, restart]Deal with form failure immediatelyAgreed disruption procedureNamed backup decision-makerCan someone stop the test during an absence?
Decision rule[Benefit, uncertainty, guardrails, decision date]Less favourable CPA interval bound no higher than −10%; guardrails passedBusiness approval ruleManagementNo change of criteria after seeing results?
Close-out[Actual period, maturity, costs, effect, deviations, decision, follow-up]Completed findings and chosen implementation routeReport export and decision recordAnalytics and managementApproval and follow-up date recorded?

Also record the time zone, currency, attribution and fully matured reference period. An account link is not a substitute for these details: settings and reports can change later. Record who may approve an extension and which observation would require a new experiment instead of a continuation.

5. Use the same measurement on both sides

In the example, a qualified enquiry must mean the same thing before and during the test. If the original variant remains subject to strict checks while the shorter form reports unchecked enquiries as successes, you are comparing two definitions. A larger number would not demonstrate an improvement in lead generation.

The reporting metric and bidding goal are separate settings

Highlighting a success metric in an experiment does not automatically determine which conversion actions are used for bidding. Primary actions in a selected standard goal normally contribute to “Conversions” and bidding; secondary actions are generally used for observation. However, selected custom goals can also use actions marked as secondary for bidding. Check these rules for primary and secondary conversion actions against the campaign's actual goal.

For setup, consult the separate guide to Google Ads conversion goals and how they are used. Keep the action, counting method, values and business qualification consistent throughout the experiment. Test both form routes through to the CRM and the intended feedback to Ads. Error pages, duplicate submissions and missing identifiers also belong in this functional check.

Agree how long sales will have to classify enquiries. If enquiries from one variant happen to arrive more often on Fridays, an early Monday analysis can underestimate their quality. Compare mature results for the same interaction periods and document unresolved cases rather than silently classifying them as unsuitable.

6. Decide assignment and the split before launch

For custom Search experiments, Google recommends a 50/50 split. Under comparable conditions, this provides substantial information for both arms. The setup instructions also distinguish between cookie-based and search-based assignment.

Repeat interactions affect the choice of method

With cookie-based assignment, a user is assigned to a variant using the cookie. This can help when repeat visits to different forms would blur the question. It does not guarantee a consistent identity across devices, browsers and consent states. With search-based assignment, each search is assigned afresh, so the same person may experience both variants. This difference belongs in the experiment question, not just the discussion of the results.

Nor does 50/50 guarantee identical daily spending, impressions or conversions. Assignment takes place at the level of eligible auctions before further targeting. Variants can develop different delivery and costs. Do not independently pause the more expensive arm simply to make the accounts look symmetrical.

Document traffic allocation and budget settings separately. Investigate unusual differences for approval issues, restrictions, bidding and tracking. An explainable difference may be part of the effect being studied; a technical outage may damage the comparison. Where users return repeatedly, also consider overlaps with other changes running in the account or on the website.

7. A change freeze protects comparability

A running experiment needs stable conditions. Adding keywords, changing ads and opening new target areas while testing a form makes interpretation unclear. Necessary maintenance remains possible, but it must be defined, traceable and assessed for its effect on both arms.

Sync does not replace your own change log

The experiment sync feature copies changes from the original campaign to the trial campaign. Sync is enabled by default, and its setting is fixed when the experiment is created. Changes in the trial do not flow back to the original. Sync transfers changes; it does not guarantee that every setting is completely identical.

Google also notes that synced changes are not identified as such in change history and that a sync error can stop subsequent changes from being copied. Keep your own log of the date, owner, reason, affected setting and actual state of each arm.

Ads may be shared between variants. Editing the control ad or making certain bulk changes can affect both sides even with sync disabled. Check how the resources are actually shared before treating “sync off” as complete separation.

In the Rheinbogen example, both form versions are therefore technically tested and clearly versioned. The offer, price, qualification rule and sales process remain fixed. If the business has to change one of these substantially, the responsible person decides whether to stop or redesign the experiment. An unnoticed mixture of changes should not continue until the planned end date.

8. Separate the primary metric, guardrails and diagnostics

The primary metric answers the original question. Guardrail metrics limit harm. Diagnostic metrics explain anomalies and suggest future hypotheses. These roles prevent you from retrospectively declaring the only positive number to have been the real objective all along.

A concise measurement agreement beats twenty success criteria

RoleMetricDefinitionSourceExample ruleTypical mistake
Primary metricQualified CPAAdvertising costs / qualified enquiriesAds using the agreed CRM actionAt least 10% improvementCounting all forms instead of qualified enquiries
GuardrailQualified volumeMature enquiries in each comparison armAds reconciled with CRMStay within the agreed volume-loss limitCelebrating a low CPA despite a steep fall in volume
GuardrailAbility to process enquiriesShare of enquiries that cannot be meaningfully handled by salesSame CRM assessment in both armsEnter the permitted rate before launchCounting pending assessments as poor leads
GuardrailExperiment costsSpending across both arms during the agreed periodAccount and budget approvalEnter the euro limit and authority to interveneTreating a budget overrun as a statistical question
DiagnosticClicks and form startsDefined interactions, examined separatelyAds and website analyticsExplain findings; do not select a winnerConfusing more starts with more suitable customers

The blank guardrail values are mandatory fields in your own protocol. Set the numbers, minimum data maturity and responsible people before spending begins. A sales team with limited capacity needs different limits from a business deliberately buying additional qualified demand.

For qualified CPA, “Conversions” in the comparison must actually represent the agreed qualified action. A blended cost-per-conversion figure covering forms, phone clicks and qualified enquiries does not meet that agreement. If qualification exists only outside Google Ads, the experiment report does not automatically calculate an interval for your external CRM metric. That requires a separate analysis using an appropriate method.

If you are testing a value-based bidding strategy instead, the conversion stage and value model must stay fixed during the strategy comparison. The previous guide to calculating reliable lead values covers this preparation. Replacing the value model and bidding strategy at the same time would create a different question with effects that are harder to attribute.

9. Check whether your data volume can support the question

A small campaign may be technically able to launch an experiment and still provide too little information for the intended decision. Other things being equal, distinguishing a ten per cent improvement is harder than detecting a very large effect. The information required also depends on variability, dependencies, assignment and the metric.

Business relevance is not a statistical minimum

Compare mature baseline volume, expected volume per arm and the maximum affordable duration. If only a few qualified enquiries arrive each week, the plan must not pretend that a small CPA difference will reliably become clear after a short period. A larger, well-founded change or an experiment at a later date may be more useful.

Campaign Guidance estimates statistical power using historical data, an assumed effect and duration. Google describes the feature for Performance Max experiments and broad match experiments in Search. It is therefore not automatically available for every custom Search test.

Power is the chance of detecting an assumed effect with the chosen method. It is neither the probability of a profitable rollout nor the confidence level of the eventual interval. An external planning calculation must state its assumptions; a simple calculation using two conversion rates is not automatically appropriate for CPA, ROAS or repeated user interactions. Skip an experiment if its realistic data budget cannot answer the actual question.

10. Plan ramp-up, measurement and data maturity separately

There is no credible universal answer of “Every Google Ads test takes 14 days.” The calendar depends on the experiment type, change, volume, weekly patterns and delays in results. Ad approval, a bidding strategy's adjustment and the later qualification of a lead are different processes.

The specific VBB timetable is a separate use case

For value-based bidding experiments , Google requires conversion values that are already being measured: at least two different, non-zero values. Google also recommends at least 50 campaign conversions in the 30 days before the split. This volume recommendation is neither a general eligibility threshold for VBB nor a guarantee of sufficient statistical power.

  1. Allow a ramp-up of two weeks or one to two conversion cycles, whichever is longer; exclude this period from the analysis.
  2. Then run the experiment uninterrupted for at least 30 days. An early statistically significant interim result does not replace this phase.
  3. Exclude recent analysis days with fewer than 90% of conversions reported, and check any additional delay in important values.

Where a two-week ramp-up is genuinely sufficient, that means at least 14 plus 30 days, followed by whatever maturation is needed at the end. Longer conversion cycles shift the entire calendar. Do not apply this specific recommendation to the form experiment without checking its relevance.

To plan your own calendar, examine the delay before conversions are reported. Costs may already be complete while results are still arriving. A historically typical lag helps with planning; it does not guarantee completeness. The conversion window, actual time to the outcome and upload delay remain separate details.

11. Monitor operations without choosing a winner every day

Daily checks are useful for detecting broken forms, disapproved ads, interrupted feedback or unacceptable spending. Declaring the best interim result to be the final outcome every day is a different action. Separate operational monitoring from the effectiveness decision agreed in advance.

Define permitted interventions before the first results

The log records the disruption, time, the two arms affected and the action taken. A protective stop does not require proof of statistical significance. The close-out might then say “Stopped because of form failure.” Without further evidence, it must not be relabelled “The variant is statistically worse.”

According to the experiment FAQs , the traffic split cannot simply be changed after setup; a different allocation requires a new experiment. Also avoid concurrent changes affecting the same campaigns or traffic. A new split midway through the comparison does not preserve the original test while merely making the budget more convenient.

Agree a maximum duration and an objective rule for extensions. “We will wait until something becomes significant” is not such a rule. If a documented interruption prevents the planned volume from arriving, first check interpretability and the time remaining. A holiday period or new offer may already have changed the comparison conditions you originally intended.

12. Read the effect, interval and benefit together

Start with the actual comparison period, data maturity, absolute figures and relative change. Add the interval and its chosen level. The current monitoring instructions describe selectable confidence intervals, with 80% as the default. Save the setting your actual report uses.

Three possible findings imply three different decisions

The following table contains freely chosen alternative teaching scenarios for Rheinbogen. The intervals are assumed report results at a 95% confidence level chosen in advance, not Google results calculated from any conversion totals shown here. “Conversions” includes only the agreed qualified action; only then does the cost-per-conversion metric used correspond to qualified CPA. A negative CPA change means an improvement. The euro figures merely apply the point estimate to the fixed €200 baseline.

ScenarioCPA effectAssumed intervalCalculated CPA10% thresholdInterpretation
A: detectable but small−3%−5% to −1%€194Even the most favourable bound falls shortDo not adopt under this business rule
B: inconclusive−10%−24% to +8%€180Point estimate reaches it; deterioration remains possibleNo reliably demonstrated economic advantage
C: more convincing−18%−24% to −12%€164Even the smallest improvement in the interval exceeds 10%Adoption possible if the other checks pass

Requiring every effect within the interval to meet or exceed the required 10% CPA reduction is a conservative business approval rule in this example. Google does not prescribe it universally. A different risk tolerance is possible, but must be documented before the test. Also check costs, qualified volume and the ability to process enquiries; no row in this table replaces those guardrail checks.

13. Do not confuse uncertainty with equality

An interval containing both negative and positive CPA effects may allow for a meaningful improvement as well as harm. It does not mean that the variants are equally good. “Inconclusive” is the appropriate finding when the evidence does not narrow down the business choice sufficiently. Demonstrating equivalence would require a specifically planned test with limits defined in advance.

Display the interval level alongside the result

The Google page on statistical methodology still describes 95% intervals, while the current monitoring guidance describes selectable levels with 80% as the default. Do not state broadly that “Google always tests at 95%.” Record the experiment type, actual chosen level and report used.

A confidence interval shows a range of effects compatible with the data and method. It does not mean “There is a 95% probability that this variant wins.” Lowering the level after seeing an inconclusive result adds no information; it changes the standard for your decision.

Google describes a method using grouped data and a jackknife procedure to estimate uncertainty. The grouped observations required for that method are not available in ordinary campaign totals. A test you calculate from two conversion counts therefore does not reproduce the official experiment report. Use additional calculations only with a clearly stated method of your own and appropriate assumptions.

14. Make an economic decision

Statistical detectability and economic benefit answer different questions. A very small, clearly detectable CPA improvement may earn less than the ongoing cost of maintaining the new solution. Conversely, a large point estimate can look attractive while its wide interval still allows for an unacceptable disadvantage.

Include effort and side effects in the same decision

At Rheinbogen, sales work is part of approval too. Making two fields optional may ease the enquiry process while also creating a need for follow-up questions. Assess this additional work using the same definition in both arms. Do not replace the agreed primary metric afterwards; check the economic limits set beforehand.

Likewise, higher total value with worse ROAS is not automatically good or bad. The business must decide in advance how much additional demand it will accept and at what efficiency. Examine absolute figures and the appropriate ratio together; an attractive ratio accompanied by a sharp contraction in business may miss the real objective.

A variant comparison also does not establish how many customers would not have been acquired without any advertising. That different question requires an appropriate incrementality design. Google describes Conversion Lift using a corresponding control group without the advertising exposure being studied. Availability must be checked separately. A Search strategy test does not gain that explanatory power through a more favourable result colour.

15. Distinguish a negative experiment from an invalid one

A properly measured deterioration teaches you about the change being tested. A broken form first teaches you about a technical failure. Confusing the two may cause you to reject a useful idea or adopt a variant on the basis of a damaged comparison.

Four disruptions need clear procedures in advance

Measurement interrupted

A conversion action or CRM feedback is missing. Mark the period, fix the cause and assess whether a complete, comparable reconstruction is possible.

Variant unavailable

A form, landing page or approval issue blocks one arm. Activate protective measures; do not treat the outage as a demonstrated effect of the planned content.

Definition changed

Qualification or values are determined differently during the test. The original question may no longer be the same; consider restarting instead of silently continuing.

Guardrail breached

Spending or operational harm exceeds the approved limit. Intervene even without significance, and record the actual reason for stopping.

Check the foundations of reliable measurement and consent before launch. Uneven measurement failures across the variants can skew the report. Consent requirements must not be bypassed to improve an experiment result.

Do not remove an unfavourable week after seeing the result. Documented exclusions require an objective reason, the same rule for both arms and an assessment of whether the remaining experiment still supports the original question.

16. Plan implementation as a separate step

An approved result does not implement itself correctly. First export the report, comparison data, settings and log. Before the experiment begins, check whether the chosen experiment type offers automatic application. Its actual setting must match the agreed business decision: if manual approval is required, automatic application must not pre-empt it.

Apply to the original or create a new campaign?

The instructions for applying a custom experiment distinguish two routes. Applying to the original ends the experiment and transfers the tested changes. Converting to a new campaign turns the trial into a regular campaign and pauses the original; the resulting campaign receives the original's budget and dates. Both performance histories are retained.

StepApply to the originalConvert to a new campaignEvidenceFinal check
ApprovalName the changes explicitlyIdentify the future campaignSigned decision recordBenefit and guardrail checks passed?
ImplementationExperiment ends; original receives changesTrial becomes regular; original pausesState after applicationCorrect campaign active?
SettingsCheck goals, assets and budgetCheck inherited budgets and datesSaved comparison of settingsNo unintended additional change?
Follow-upMonitor delivery and guardrailsMonitor delivery and guardrailsDate and responsible personFull delivery operationally sustainable?

Follow-up remains necessary because full delivery and a later market period bring different conditions. It is operational monitoring after the experiment, not new randomised proof. Avoid introducing another major change at the same time if you first want to assess whether the approved version has been implemented correctly.

17. Frequently asked questions about Google Ads experiments

How long must a Google Ads experiment run?

There is no single suitable duration for every type. Plan around the specific change, volume, conversion cycles and required precision. Separate ramp-up, the period suitable for analysis and data maturity. The VBB use case described here has its own recommendations: exclude ramp-up, then run for at least 30 days. A calendar date alone proves neither that there is enough information nor that all results have arrived.

Do I need at least 50 conversions for every test?

No. The recommendation cited here concerns Google's VBB campaign experiments: at least 50 campaign conversions in the 30 days before the split. It is not a universal minimum for every Search question, nor a guarantee of sufficient power. The commercially important effect size, variability and assignment still matter. At low volume, skipping the test or asking a different question may be more useful than running a comparison that is barely interpretable.

Why do two 50/50 arms spend different amounts?

The split does not promise equal spending. Assignment, targeting, auction participation and resulting outcomes are different steps. Investigate unusual differences for approval issues, restrictions and technical problems before calling them errors. Do not force identical totals through unplanned interventions. With other splits, also check whether the report already normalises results; normalising them again could distort the comparison.

Should I use cookie-based or search-based assignment?

The decision depends partly on whether repeated exposure to both variants by the same person would interfere with your question. Cookie-based assignment supports more consistent allocation using the cookie, but does not guarantee identity across devices. Search-based assignment allocates each search afresh, so returning searchers may experience both variants. Document the chosen method and its limitations before launch, particularly where buying decisions take longer.

Does a significant result automatically mean I should adopt the variant?

No. The effect may be detectable yet too small to matter commercially. Data maturity, comparability and guardrail metrics must also be sound. Record the confidence level actually selected; the current monitoring guidance describes selectable intervals with 80% as the default. A marker without a level and business rule is not a complete basis for a decision. Also examine the work and risks involved in permanent implementation.

What if there is no winner after the planned duration?

First archive an inconclusive finding. Check which meaningful benefits and disadvantages the interval still allows. An extension must fit an objective framework agreed beforehand; running indefinitely until the desired colour appears is not a reliable rule. You may need a more distinct variant, more usable volume or a newly planned experiment. A lack of significance does not demonstrate equivalence.

May I change ads or goals during the experiment?

These changes can undermine the original question and affect both arms through sync or shared ads. Define permitted maintenance in advance and keep your own log. A necessary safety fix is possible; afterwards, interpretability must be reassessed. Changing the conversion stage or value definition, in particular, can turn a pure strategy test into a different comparison.

Does the test prove that Google Ads brings additional customers?

A variant experiment examines a change against another advertising configuration. It does not automatically answer how much additional business occurs compared with no exposure to the advertising being studied. That requires an appropriate incrementality design. Product availability and the measurement method must be checked separately. In the final record, describe the decision actually tested and avoid broader revenue promises.

18. Close every experiment with a traceable decision record

A good close-out contains more than “B wins.” Record the actual comparison dates, excluded phases, maturity, primary metric, effect and interval level. Add guardrail metrics, deviations and the economic reason for your decision. This keeps the finding understandable when the account looks different later.

Turn a test into a justified next action

Rheinbogen's record might read: “The form variant will be adopted because the fully checked comparison meets our defined benefit rule and all guardrails. Implementation will be reviewed on the named date.” A different finding should state just as clearly “Inconclusive” or “Uninterpretable because of measurement failure,” together with the appropriate next action.

Keep the hypothesis and result together. Diagnostic findings may suggest new questions, but must not retrospectively become the confirmed original hypothesis. Schedule the next overlapping experiment in a sensible sequence. An experiment programme becomes valuable through repeatable decisions, not through the number of tests running simultaneously.

Would you like to establish which change your account can reliably test with its actual volume? Salestudia's Google Ads management service connects measurement checks, experiment planning and economic analysis to support a decision you can justify.