
The dashboard says scale. The business needs a different answer.
You are looking at an ad account that appears to be working. Reported conversions are steady or rising. The platform is recommending more budget. The modeled conversion count fills in gaps caused by missing identifiers, consent restrictions, cross-device behavior or other measurement problems.
The sensible answer is not to ignore the platform. It is to separate two decisions that often get treated as one.
Use the platform’s reported and modeled conversions to help it optimize delivery. Use an incrementality test before making a material budget increase when the cost of being wrong matters. The first question is, “Which users should receive the ad?” The second is, “How many of these outcomes would have happened without the ad?” Modeled attribution can help with the first. It does not answer the second.
That is why my answer is usually yes, run an incrementality test before scaling meaningfully. Do not turn that into a rule that every small budget adjustment needs a formal study. A modest, reversible increase can be treated as an operating test. A large increase, a new channel, a move into broad automation, or a decision based mainly on modeled conversions deserves causal evidence. This is the same guardrail logic discussed in AI Can Change Ad Budgets. It Should Not Change Them Without Guardrails.
Modeled conversions improve measurement. They do not prove causation.
Google’s own explanation is more careful than many ad reports. A modeled conversion is usually an estimate of the missing link between an ad interaction and a conversion that occurred. Google says the model is not determining whether the conversion happened. It is estimating whether the ad interaction led to it. (support.google.com)
That distinction matters. Suppose a homeowner was already searching for your company, clicked a branded ad, and booked an appointment. The platform may reasonably assign credit to that interaction. But the commercial question is whether the ad created an additional appointment, changed the timing, or simply intercepted demand that was already on its way.
Google says it uses holdback validation and only includes modeled conversions when it has high confidence in the result. That is useful quality control, but it is still validation of the model’s attribution estimate. It is not the same as withholding advertising from a comparable group and observing the difference in outcomes. (support.google.com)
This is not an argument that modeled conversions are fraudulent or worthless. They can make bidding less biased than using only the conversions the platform happened to observe. They are useful for day-to-day optimization. The mistake is promoting that operational signal into a financial claim: “If we double spend, we will double these customers.”

Which measurement should control the decision?
| Criterion | Platform attribution | Incrementality testing |
|---|---|---|
| Best use | Daily bidding, creative and audience optimization | Budget increases, channel validation and stop-or-scale decisions |
| What it measures | Conversions the platform can observe or model and associate with ads | The difference between outcomes with advertising and outcomes without it (better) |
| Speed | Fast, often available while campaigns run (better) | Slower because a credible control group and test window are required |
| Small-account practicality | Available to almost every advertiser (better) | Can be noisy when conversion volume or geographic coverage is low |
| Protection against demand capture | Limited, because the platform evaluates its own attributed outcomes | Strongest when treatment and control are well designed (better) |
The practical threshold is not “small business” versus “large business.”
The better question is whether the proposed increase is material enough to justify learning what the next dollars actually buy.
A local service company spending a few thousand dollars a month may not have enough volume for a clean user-level lift study. A geo holdout may still be possible if it serves several comparable markets. An ecommerce shop with one national market may need a platform experiment, a budget-step test, or a broader business metric such as blended customer acquisition cost and total revenue.
Practitioners working with smaller accounts report three workable designs: pause branded activity in a defined market or period, hold out matched geographies, or increase budget in a planned step and compare the marginal cost of the added conversions with the old baseline. One practitioner guide positions these tests for accounts spending between $2,000 and $10,000 per month, with a two-week brand pause, a four-week geo split and a roughly seven-week budget-step design. Those are field recommendations, not universal thresholds. (entropyand.co)
That qualification matters because a test can fail in two opposite ways. It can be too small to detect the lift that would change your decision. Or it can be so disruptive that the business loses more money holding out spend than it could reasonably learn. A practitioner guide to geo tests makes the same point: market selection, seasonality, baseline noise and statistical power all interact. A weak holdout can produce a precise-looking number that should not be trusted. (semdispatch.com)
For a small business, the right test is therefore the cheapest design that can answer the actual decision. If the decision is whether to add $1,000 next month, you do not need an academic measurement program. You do need a pre-agreed rule for what result would justify that extra spend.

A useful pre-scale test has four decisions
-
Choose the business outcome
Use booked jobs, completed purchases, qualified leads or profit contribution, not the platform’s preferred proxy unless that proxy is the real commercial goal.
-
Define the counterfactual
Decide what will represent “without the ads”: a randomized audience holdout, matched geographic control, brand pause or planned budget baseline.
-
Keep the test narrow
Change one important variable. Do not change budget, landing page, offer, creative and sales process at the same time.
-
Set the scale rule before launch
Write down the minimum lift, incremental cost per acquisition or incremental return that would make the increase worthwhile.
Use the platform’s experiments where they can answer the question
The platforms are not blind to this problem. Google describes Conversion Lift as a controlled experiment that separates treatment and control groups and calculates the difference in downstream conversions. It can report incremental conversions, incremental cost per action and incremental return on ad spend when the account has the required data. Google also offers geography-based studies that can incorporate offline outcomes. (support.google.com)
Google’s Performance Max experiments include uplift tests and upgrade tests. Those are directly relevant when the decision is whether to add Performance Max or shift budget from an existing campaign. (support.google.com) But availability and volume requirements matter. Conversion Lift is not available to every account, and Google’s custom experiment guidance says reliable results require sufficient conversion volume, with one current example being more than 100 daily conversions for the base campaign. That is a platform requirement for reliable results in that experiment context, not a universal minimum for every incrementality method. (support.google.com)
If you cannot run a platform lift study, do not pretend the answer is impossible. Use a geo design when markets are genuinely comparable. Use a budget-step design when you cannot turn a channel off. Compare total business outcomes, not just the campaign’s reported conversions. A rise in paid conversions alongside flat total sales is a warning sign. So is a strong platform ROAS paired with worsening blended CAC.
The most useful result may not be “ads work” or “ads do not work.” It may be a calibration factor. If the platform reports 100 conversions and a credible test suggests only part of that volume is incremental, you can use the finding to improve future budget decisions while continuing to use attribution for campaign management. Google’s broader measurement guidance recommends this combined approach: attribution, incrementality and media mix modeling each answer different questions. (thinkwithgoogle.com)
Numbers worth putting beside the decision
Google says modeled online conversion data can take up to 5 days to fully process and stabilize. Source: About modeled online conversions.
Google gives more than 100 daily conversions as an example minimum requirement for reliable results in custom experiments. Source: About custom experiments.
A practitioner guide describes small-business PPC incrementality tests for accounts spending between $2,000 and $10,000 per month. Source: Incrementality Testing for Small Business PPC.
A practitioner guide describes small-business PPC incrementality tests for accounts spending between $2,000 and $10,000 per month. Source: Incrementality Testing for Small Business PPC.
The practitioner example compares a $55 baseline CPL with a $180 marginal CPL after a budget step. Source: Incrementality Testing for Small Business PPC.
Scale after evidence, not because the dashboard became more confident
A small business should not wait for perfect measurement before advertising. It should avoid making an irreversible budget commitment on a number that answers a different question.
Modeled conversions can make an ad platform’s optimization better informed. Incrementality testing tells you whether the advertising changed the business outcome. Use the first to run campaigns. Use the second when the next budget decision is large enough that false confidence would be expensive.
In practice, that means a modest increase can be staged and monitored, while a major increase should have a pre-scale test or a clearly documented reason why testing is not feasible. The important discipline is not statistical ceremony. It is refusing to confuse “the platform found a conversion” with “the next dollar created a customer.”