Written by: Aaron Rovner, Founder, Saas Hero | Last updated: August 27, 2026
Key Takeaways
- Platform ROAS overstates true performance by 2–5x because it relies on attribution instead of incremental revenue from controlled experiments.
- The 0–100 scorecard ranks agencies using five weighted components: verified incremental revenue (30%), measurement confidence (25%), client retention (20%), gross-profit contribution (15%), and spend-tier normalization (10%).
- Evidence quality tiers A–E set scoring ceilings. Only geo holdouts or matched-market tests with CRM closed-won data at 95%+ confidence qualify for Tier A and a top score.
- Spend-tier normalization adjusts results for auction difficulty so $15k/month and $150k+/month accounts can be compared fairly.
- Book a discovery call with SaaSHero to get the working scorecard template and apply it to your agency shortlist.
1. ROAS Overstatement and the Incrementality Gap
Most agency shortlists include at least one firm reporting a 4x or 5x platform ROAS. The incrementality gap is the difference between that number and what a controlled experiment would show. For B2B SaaS, this gap is structural, not accidental.
A common real-world gap is 40–60% overstatement when companies reconcile platform ROAS against CRM-sourced closed-won revenue. A reported 4x ROAS may reflect a true 2x–2.5x when measured against booked revenue. Three compounding factors drive the overstatement: last-click attribution, proxy conversion events such as form fills instead of closed-won deals, and data-driven attribution models trained on click patterns that favor bottom-of-funnel credit.
B2B SaaS sales cycles have lengthened 22% since 2022 (Optifai Pipeline Study, 2026, N=939). That shift widens the gap between Google’s 30-day attribution window and actual revenue realization.
A holdout test example illustrates the gap. Platform-reported revenue can exceed truly incremental revenue, which yields a higher platform ROAS than the true iROAS.
Use the following steps to detect the gap before committing to an agency:
- Pull closed-won revenue by campaign from the CRM for the last 90 days and divide by the ad spend reported in Google Ads for the same period. Compare that figure to the platform ROAS the agency reports.
- Ask the agency whether any holdout or geo-lift test has been run on the account. Request the test design, the control group definition, and the statistical confidence level.
- Check whether the primary conversion action feeding Smart Bidding is a CRM-qualified event (opportunity created, SQL, closed-won) or a proxy event (form fill, page visit, content download).
- Request the search terms report for the last 30 days and identify what share of spend went to non-ICP queries.
Pitfall: Agencies sometimes present conversion lift studies run inside Google Ads as incrementality evidence. Platform-native lift studies measure lift within the platform’s own attribution model and cannot correct for double-counting across channels or for organic demand the platform would have captured anyway. Only geo holdouts or matched-market tests using backend CRM data qualify as Tier A or B evidence under the scoring model below.
2. The 0-100 Scoring Model with Exact Weights and Formula
The scorecard creates a single comparable score so agencies compete on revenue impact instead of on reporting tricks. Because a $15k/month account and a $150k/month account operate in structurally different auctions, Component 5 normalizes results for spend tier to ensure fair comparison. The mechanics of that adjustment appear in Section 4.
The formula is: Total Score = (Component 1 × 0.30) + (Component 2 × 0.25) + (Component 3 × 0.20) + (Component 4 × 0.15) + (Component 5 × 0.10). Each component is scored 0–100 before weighting.
Component definitions and scoring decision criteria:
- Component 1 — Verified Incremental Revenue (30 points): Score 80–100 for iROAS verified by a geo holdout or matched-market test using CRM closed-won data at 95%+ confidence. Score 50–79 for iROAS estimated from offline conversion import with opportunity-stage weighting. Score 0–49 for platform ROAS only, with no CRM reconciliation.
- Component 2 — Measurement Confidence (25 points): Score 80–100 for a fully documented primary and secondary conversion architecture with lifecycle-stage events flowing back to the ad platform from the CRM. Score 50–79 for offline conversion import configured but not validated against closed-won data. Score 0–49 for form-fill-only conversion tracking with no CRM connection.
- Component 3 — Client Retention (20 points): Score 80–100 for 85%+ year-over-year retention across the agency’s B2B SaaS book. Score 50–79 for retention at or above the median for B2B marketing agencies. Score 0–49 for retention below the median or undisclosed.
- Component 4 — Gross-Profit Contribution (15 points): Score 80–100 for pipeline and closed-won revenue reported net of media spend with CAC payback documented. Score 50–79 for pipeline reported without payback period. Score 0–49 for lead volume or CPL only.
- Component 5 — Spend-Tier Normalization (10 points): Applied via the table in Section 4. This component adjusts raw scores for the spend level at which results were achieved.
Use neutral benchmark references for calibration. Top-quartile B2B agency specialists hold 85% or more client retention. Median iROAS across geo tests is 2.3x. A healthy LTV:CAC ratio for SaaS is generally 3:1, and CAC payback under 12 months is considered strong.
3. Evidence Quality Tiers A–E with Scoring Rules and Examples
The 0-100 score is only as reliable as the evidence behind each component. Evidence tiers A through E classify the data quality an agency can produce for any claim about revenue impact. Tiers determine the maximum score available in Components 1 and 2. An agency presenting Tier D evidence cannot score above 49 on verified incremental revenue regardless of the number it reports.
Tier definitions, scoring rules, and anonymized examples:
- Tier A — Geo holdout or matched-market test, CRM closed-won data, 95%+ confidence: Maximum score 100. Example: a B2B SaaS agency paused paid search in three matched DMAs for eight weeks, compared closed-won ARR from CRM against the control markets, and produced an iROAS of 2.8x at 97% confidence. Geo holdout results should be validated with high statistical confidence before large-scale budget reallocations.
- Tier B — Offline conversion import with opportunity-stage weighting, CRM-reconciled: Maximum score 79. Example: an agency imports SQL-created and closed-won events from HubSpot into Google Ads as primary conversions, reconciles platform-reported pipeline against CRM pipeline monthly, and documents a 15% variance. The practical fix for ROAS overstatement is to import offline conversion data from CRM systems, specifically opportunity creation and closed-won events with weighted values.
- Tier C — Multi-touch CRM attribution (W-shaped or full-path), no holdout: Maximum score 64. Example: an agency uses HubSpot’s W-shaped model across closed-won deals and reports influenced revenue by campaign. HubSpot attribution reports show influenced revenue rather than caused revenue. A channel credited with $150,000 means it appeared in buyer journeys of customers who closed that amount, not that it directly caused the revenue.
- Tier D — Last-touch or platform-native attribution only: Maximum score 49. Example: an agency reports Google Ads conversion value using last-click attribution with form fills as the primary conversion action. As noted in Section 1, platform ROAS reflects attribution, not causation, so agencies presenting only last-click data cannot score above 49 in Component 1.
- Tier E — Self-reported or unverified claims, no CRM connection: Maximum score 20. Example: an agency presents a case study citing “4x ROAS” with no methodology, no CRM data, and no client reference available for verification.
Because only Tier A and B evidence can unlock the top scoring bands in Components 1 and 2, the following requirements define what qualifies an agency’s data package for those tiers.
Evidence requirements for Tier A or B classification:
- Test design documented with control group definition, market selection rationale, and run duration of at least 6–12 weeks for B2B categories to capture the full purchase cycle
- Revenue source is CRM backend data (Salesforce or HubSpot closed-won), not platform-reported conversions
- Statistical confidence reported at 95% or above with standard error documented
- Results reconciled against total account revenue, not platform attribution alone
- Client reference available to confirm the test was run as described
Pitfall: Salesforce Campaign Influence attribution experiences silent failures when Contact Roles are missing on Opportunities. An agency presenting Salesforce-sourced revenue data without confirming Contact Role hygiene may be presenting Tier D evidence labeled as Tier C.
4. Spend-Tier Normalization and the Final Headline Metric
A $15,000/month account and a $150,000/month account operate in structurally different auctions. Comparing raw iROAS figures across those spend levels without normalization rewards agencies that inherited high-intent, low-competition accounts instead of those that produced genuine incremental lift. Spend-tier normalization adjusts Component 5 scores to account for the difficulty of the environment in which results were achieved.
| Monthly Ad Spend Tier | Median Platform ROAS Benchmark | Median iROAS Benchmark | Normalization Multiplier for Component 5 |
|---|---|---|---|
| $15k–$30k | ~1.55x average platform ROAS for B2B SaaS on Google Ads (Gawa Growth analysis, not attributed to Varos) | 2.3x (baseline) | 1.00 (baseline) |
| $30k–$75k | ~1.55x average platform ROAS for B2B SaaS on Google Ads (Gawa Growth analysis, not attributed to Varos) | 2.3x (baseline) | 1.10 (moderate auction saturation) |
| $75k–$150k+ | ~1.55x average platform ROAS for B2B SaaS on Google Ads (Gawa Growth analysis, not attributed to Varos) | 2.3x (baseline) | 1.20 (high saturation, diminishing marginal return expected) |
The final headline metric produced by the scorecard is Verified Incremental Revenue per $1 Ad Spend. Calculate it as (CRM closed-won revenue attributable to the campaign via holdout or offline import) ÷ (total ad spend in the measurement period). This figure is spend-tier normalized by multiplying the raw Component 5 score by the appropriate multiplier before applying the 10% weight. That number gives a CFO or PE operating partner a single metric to compare agencies across a portfolio without a methodology debate.
Use these steps to apply the model:
- Collect the agency’s evidence package for each component and assign a Tier A–E classification before scoring.
- Score each component 0–100 based on the tier ceiling and the decision criteria in Section 2.
- Apply the spend-tier multiplier to Component 5 before weighting.
- Multiply each component score by its weight and sum to produce the 0–100 total.
- Divide the agency’s verified incremental revenue figure by its total ad spend in the same period to produce the headline metric for board reporting.
Pitfall: Google non-brand search can report more revenue than it truly delivers, with platform-reported ROAS exceeding the true incremental ROAS in incrementality testing. An agency presenting non-brand search results without a holdout test is likely presenting a figure above the verified incremental number. Apply the Tier D ceiling (49 points maximum) to any Component 1 score built on platform-reported non-brand search ROAS alone.
Frequently Asked Questions
Incremental ROAS vs Platform ROAS for Agency Evaluation
Platform ROAS is calculated by dividing the revenue Google’s attribution model credits to your ads by the amount you spent. It is a correlational metric. It tells you that revenue and ad spend occurred together, not that the ads caused the revenue.
Incremental ROAS (iROAS) divides only the revenue that would not have existed without the ads by the same spend figure. Controlled experiments verify that incremental portion. Platform ROAS can be inflated by 30–60% or more through last-click attribution, proxy conversion events, and organic demand the platform claims credit for.
An agency evaluated on platform ROAS alone can report strong performance while the CRM shows flat pipeline. iROAS, anchored in closed-won CRM data, removes that ambiguity and gives boards a defensible number tied to actual revenue.
Timeline for Implementing the 0-100 Scorecard
You can apply the scoring framework to a shortlist of two to four agencies in one to two weeks if each agency supplies its evidence package promptly. The most time-consuming step is requesting and verifying the evidence behind Component 1 (verified incremental revenue) and Component 2 (measurement confidence). Agencies presenting Tier A or B evidence need to supply test designs, CRM data exports, and statistical confidence documentation instead of a slide deck.
If an agency cannot produce that documentation within five business days, that response is itself a scoring signal. Assign Tier E and cap the Component 1 score at 20. For organizations that have not yet run a holdout test on their own account, SaaSHero’s CRM-connected stack can establish the measurement baseline during onboarding. That setup positions the first 90-day validation period as the initial evidence-collection window.
Internal Ownership of the Scorecard Process
The VP of Marketing or CMO owns the final score and the agency selection decision. RevOps or Marketing Operations should validate the evidence tiers for Components 1 and 2. That work includes assessing CRM data quality, Contact Role completion rates, offline conversion import configuration, and lifecycle-stage definitions.
Those checks require access to the CRM and tag management systems that the marketing leader typically does not operate directly. The CFO or VP of Finance should review the spend-tier normalization table and the final Verified Incremental Revenue per $1 Ad Spend metric before the decision is finalized. For PE-backed companies, the operating partner’s involvement at the evidence-tier review stage compresses the evaluation timeline and aligns the methodology with how the fund measures marketing efficiency across the portfolio.
Adapting the Scorecard for Smaller Teams or Lower Spend
The weights and formula apply at any spend level above $15,000 per month. That level is the floor at which there is sufficient conversion volume to run statistically meaningful holdout tests. Below that threshold, Tier A evidence is rarely achievable because geo holdouts require enough conversion events in both treatment and control markets to detect a 5–10% minimum effect at 80% statistical power.
For teams at the lower end of the $15k–$30k tier, Tier B evidence (offline conversion import with opportunity-stage weighting) is the realistic ceiling for Component 1. The scorecard should be scored with that ceiling in mind. Smaller internal marketing teams, such as the two-to-four-person functions common at $10M–$50M B2B SaaS companies, can run the scorecard with the VP of Marketing leading the process and RevOps validating the CRM evidence.
How SaaSHero’s Measurement Stack Maps to Evidence Tiers
SaaSHero’s standard onboarding rebuilds conversion tracking from the ground up. The team establishes a primary and secondary conversion architecture where only CRM-qualified events, such as SQL creation, opportunity creation, and closed-won, feed Smart Bidding as primary conversions. Lifecycle-stage events are pushed back from the CRM into the ad platforms, and Looker Studio dashboards connect ad spend to pipeline and closed-won revenue in the client’s own CRM.
That configuration satisfies the Tier B evidence requirements for Component 1 and Component 2 from the start of an engagement. It also positions the account for Tier A evidence once a geo holdout or matched-market test is run. This setup is the only configuration in which the Verified Incremental Revenue per $1 Ad Spend headline metric can be calculated from CRM data rather than estimated from platform attribution.
Conclusion: Apply the Scorecard and Move Forward
Platform ROAS behaves as a self-fulfilling metric. It trains algorithms toward whatever conversion event it is given, reports improving efficiency as the algorithm gets better at finding that event, and can leave the CRM showing flat pipeline. The 0-100 scorecard replaces that loop with four verifiable steps.
- Assign evidence tiers A–E to each agency’s revenue claims before scoring any component.
- Score the five weighted components, including verified incremental revenue, measurement confidence, client retention, gross-profit contribution, and spend-tier normalization, using the tier ceilings as hard caps.
- Apply the spend-tier normalization multiplier to Component 5 before weighting.
- Divide verified incremental revenue by total ad spend to produce the single headline metric for board reporting.
The agencies that score highest under this framework already run CRM-connected measurement, already separate primary from secondary conversions, and already produce evidence that survives a holdout test. SaaSHero’s existing stack satisfies those requirements at onboarding. That capability is why it is often the only agency on a shortlist that can supply Tier B evidence from day one and a path to Tier A within the first engagement cycle.