Written by: Aaron Rovner, Founder, Saas Hero

Key Takeaways

  • Most agency-versus-in-house evaluations are unfair because they use identical metrics even though each party controls different variables.
  • A fair evaluation uses a shared scorecard that separates controllable from uncontrollable factors and scores both models on CRM-connected outcomes such as qualified pipeline, CAC payback, and LTV:CAC.
  • Agencies usually control ad accounts and campaign execution, while in-house teams control CRM, brand, and post-click experience, so scoring both on CPL or ROAS misleads decision-makers.
  • Cost comparisons must use total cost of ownership for both models, including fully loaded salaries, tooling, recruiting, and management overhead for in-house teams versus agency retainers.
  • SaaSHero provides an outsourced inbound growth team that owns the full chain from impression to CRM record, which removes the scope gaps that create unfair evaluations.

Get Your Fair Evaluation Scorecard

Why Most Agency-Versus-In-House Evaluations Break

Most evaluations fail because they score an agency and an in-house team on identical metrics even though each controls different variables. This creates a comparison that looks rigorous but does not reflect how work actually happens.

Agencies typically control ad accounts, campaign architecture, creative, and bidding strategy. They do not control landing pages, CRM lifecycle definitions, sales follow-up speed, or the conversion events that feed the algorithm. In-house teams control the CRM, brand voice, and post-click experience. They rarely control paid-media specialization at the depth required to manage a $15,000-per-month account across Google, LinkedIn, and Meta simultaneously.

When both models are scored on the same CPL or ROAS target, the agency gets penalized for landing page conversion rates it does not own, and the in-house team gets penalized for paid-media depth it was never hired to provide. Lara Galvin of Codify Consulting frames this precisely: holding a function accountable for outcomes it does not control produces noise, not rigor.

A shared scorecard fixes this by mapping what each party actually controls before any metric is applied. Controllable factors are scored. Uncontrollable factors such as market conditions, sales cycle length, product-market fit, and sales team follow-up speed are documented as context instead of treated as failures.

For a VP of Marketing at a $20M–$50M B2B SaaS company, this distinction separates a defensible board presentation from a political fight with both her agency and her internal team.

How To Measure Agency And In-House Performance Fairly

A fair evaluation follows five steps in sequence so each step builds on the last.

  1. Establish A Common Baseline. Define the same measurement window, attribution model, and CRM stage definitions for both models before scoring anything. A 90-day window is the minimum for B2B SaaS. The median B2B SaaS sales cycle runs 84 days, so shorter windows cannot be evaluated against revenue outcomes.
  2. Separate Controllable From Uncontrollable Factors. List what each party actually controls, such as ad accounts, landing pages, CRM, lifecycle definitions, brand, and sales follow-up, and score only what they control. An agency that does not own the landing page cannot be held accountable for post-click conversion rate.
  3. Audit Resource Utilization. Track hours spent, revision cycles, and bottleneck points. Internal approval delays are a client-side factor. Agency turnaround times are an agency-side factor. Client-side factors such as slow access approvals, delayed feedback, and unclear decision-making can cause bad agency performance and must be separated from agency execution failures before any score is assigned.
  4. Score Against Shared Business Outcomes. Use CRM-connected metrics such as qualified pipeline created, cost per sales-qualified lead, CAC payback, LTV:CAC, and pipeline coverage. Cometly’s B2B SaaS analytics framework recommends treating acquisition metrics like CPL as guardrails while pipeline and revenue metrics serve as the primary targets.
  5. Review Strategic Value. Assess whether the party acts as a tactical order-taker or brings data-driven recommendations, market trends, and scaling strategies. An agency that waits to be told what to test represents a structural issue, not a single-person problem.

How In-House And Agency Marketing Actually Differ

In B2B SaaS, in-house and agency models carry different responsibilities, cost structures, and control surfaces. They perform different jobs, so identical scorecards distort reality.

Scope: Agencies typically own paid channels such as search, social, and display. In-house teams own brand, CRM, and lifecycle. Most pipeline leaks appear in the gap between those scopes.

Control: In-house teams control the CRM and post-click experience. Agencies control ad accounts but often not landing pages, which Stackmatix identifies as the single highest-leverage variable in the conversion funnel.

Speed: Agencies launch faster because processes, tooling, and team structures already exist. Building an in-house marketing team takes 3–6 months of recruiting, onboarding, and ramp-up, while an agency can launch campaigns within weeks.

Cost Structure: Agency retainers are a fixed or spend-indexed monthly fee. In-house hires carry salary, benefits, tooling, training, management overhead, and recruiting costs. Chrysales advises budgeting 30–40% on top of base salary for total cost per employee, which turns an $80,000 salary into roughly $110,000 before software subscriptions.

Specialization: Agencies bring cross-client pattern recognition from managing spend across many accounts in the same category. In-house teams bring product and customer knowledge no agency will match. Each advantage applies to a different part of the funnel.

For a $20M–$50M B2B SaaS company, a 2–4 person in-house team covering content, product marketing, events, and lifecycle rarely includes a paid-media specialist. That is the structural gap an outsourced growth team fills. It provides the execution layer in-house teams are not staffed to deliver while leaving in-house judgment intact.

The Shared Scorecard: Which Metrics Apply To Both Models And Which Do Not

The table below separates metrics that are fair to score against both models from metrics that systematically mislead. The pattern is consistent: metrics that stop at the form fill mislead, while metrics that connect to CRM revenue create a fair comparison.

Metric Fair to Both Models? Why It Applies or Misleads
Qualified pipeline created Yes CRM-connected; measures business outcome regardless of who controls the channel.
Cost per sales-qualified lead Yes Ties spend to sales-accepted outcomes, not form fills. The median cost per SQL for B2B SaaS Google Ads runs $800–$2,500 in 2026, varying by vertical.
CAC payback Yes Under 12 months is strong for venture-backed B2B SaaS at the growth stage; comparable across models when fully loaded costs are included.
LTV:CAC Yes 3:1 is generally considered healthy for SaaS; applies to both models when CAC includes all acquisition costs such as ad spend, agency fees, salaries, and tooling.
Pipeline coverage Yes Measures whether pipeline is sufficient against a sales target. Both models can be held to this metric.
CPL No Penalizes agencies that optimize to qualified outcomes and rewards cheap, unqualified form fills. A $65 CPL at a 30% pipeline-fit rate produces a $217 pipeline-qualified CAC, while a $160 CPL at 55% fit produces $291, which makes the cheaper CPL channel worse on pipeline cost.
ROAS No Misleads in B2B SaaS with long sales cycles and buying committees. Platform-reported ROAS excludes CRM outcomes and stops at the form fill.
Impressions No Vanity metric that does not connect to pipeline or revenue.
Form-fill volume No The cheapest form fillers in B2B SaaS are typically students, job seekers, competitors, and companies below the ICP threshold, none of whom become paying customers.

The measurement gap is structural, not accidental. OneMetrik’s audit of a B2B SaaS account spending $25,000–$50,000 per month on Google Ads found that CPL improved 35% while pipeline value fell 25% in the same period, because smart bidding hit the wrong objective. Platform-reported metrics stop at the form fill. CRM-connected metrics connect to revenue, and a fair scorecard uses only those.

See How SaaSHero Scores On Your Scorecard

How To Compare Agency And In-House Costs Fairly

A fair cost comparison looks at total cost of ownership on both sides instead of comparing an agency retainer to a single in-house salary.

Agency total cost of ownership includes the retainer, any media management fees, creative production if billed separately, landing page builds, reporting tooling, onboarding time, and internal management hours spent on approvals and strategy calls. At SaaSHero, the retainer is flat and indexed to total monthly ad spend rather than channel count. That structure removes the pricing incentive to recommend more channels. Adding a channel does not raise the fee, so channel-mix recommendations are based on evidence.

In-house total cost of ownership includes base salary plus benefits, which Long Weekend’s 2026 analysis puts at 30–50% above base salary once employer taxes, health benefits, paid time off, retirement contributions, and workers’ compensation are included. Add tooling, where a typical in-house marketing software stack costs $1,000–$2,500 per month, plus training, management overhead, and recruiting costs. Recruiting an experienced paid media lead commonly costs 15–25% of first-year salary when external recruiters are used. A fully loaded in-house hire includes every one of these line items.

T.A. Monroe’s 18-month cost model for a senior B2B SaaS paid media hire estimates $257,000–$390,000 all-in excluding ad spend, while a B2B SaaS agency retainer over the same period runs $270,000–$540,000 depending on scope. The ranges overlap, so the comparison must be made on scope equivalence rather than sticker price.

A single in-house paid media hire is also a single point of failure. When that person resigns, the ramp clock restarts with account history and platform knowledge walking out the door. An agency retainer avoids that key-person risk.

Both models should be evaluated against the same thresholds, using the benchmarks in the scorecard above and including fully loaded costs in the CAC calculation on both sides.

For more on the cost structure comparison, see Cost of In-House Paid Media vs a B2B SaaS Agency. Cost is only one axis, though. The next section shows how behavior reveals when each model is failing.

Warning Signs For Each Model

Each model fails in different ways, so the red flags you watch for should match the structure you have in place.

Agency red flags:

  • Reporting that leads with impressions, reach, or form-fill volume instead of pipeline and CAC.
  • No ownership of landing pages. Agencies that do not own the post-click experience cannot be accountable for conversion outcomes.
  • Reactive behavior where the client sets the test agenda and chases status updates.
  • Stagnant campaign structure with the same keywords, audiences, and creative running for more than two quarters without a documented rationale.
  • Generic messaging that could belong to any competitor in the category.
  • Creative that does not improve, with variations instead of true tests and no evidence of learning.

In-house red flags:

  • One person covering five specializations such as paid search, paid social, creative, landing pages, and attribution, with shallow depth in each.
  • Post-click experience under-served, with traffic landing on the homepage or a product page that does not match the campaign audience.
  • Attribution plumbing neglected. Enterprise marketing teams typically operate platforms at only 30% utilization, which leaves CRM integrations and conversion tracking incomplete.
  • No paid-media specialist. The team covers content, product marketing, and lifecycle, but nobody manages the ad account at the depth that spend level requires.
  • Work queuing behind a single approver, which slows every campaign launch and creative test.

In both cases, the root cause is structural rather than personal, so the evaluation design needs to reflect that reality. See the full agency vs in-house framework for a 2026 B2B perspective.

How To Run The Evaluation Without Demoralizing Your Team Or Prematurely Firing Your Agency

A political-safety layer keeps the evaluation from turning into a blame exercise. Without it, a VP of Marketing risks demoralizing her internal team and triggering a defensive response from her agency before the data is even reviewed.

Frame the evaluation as a structural diagnostic rather than a performance review. The question to ask is “which variables does each party control, and are they being scored on those variables?” instead of “who is failing?” That framing is defensible to both parties and to a CFO.

Separate controllable from uncontrollable factors before assigning any score. If the agency’s landing page conversion rate is low but the agency does not own the landing page, move that metric off the agency’s scorecard and into the in-house column or flag it as an unowned gap.

Share the scorecard with both parties before the evaluation begins. An agency that sees the evaluation criteria in advance can prepare evidence. An in-house team that sees the criteria understands they are being assessed on what they control instead of on paid-media depth they were never hired to provide.

Use descriptive ratings such as exceeds expectations, meets expectations, and below expectations, with evidence attached. Numerical scores often create false precision and invite arguments about methodology instead of substance.

Most agencies need three to six months to ramp fully. Performance in that window should be evaluated against leading indicators such as pipeline trajectory, SQL rate, and landing page conversion rate rather than revenue results alone, because a B2B SaaS deal closing 90 days after the campaign launch will not appear in a 60-day evaluation window.

Client-side factors can also drag down agency performance. Slow access approvals, delayed creative feedback, unclear decision-making, and misalignment between what marketing is asked to generate and what sales can close are all client-side variables. Document them as context before scoring the agency on outcomes it could not control.

Run Your Next Evaluation With SaaSHero’s Framework

Where SaaSHero Fits In Your Evaluation

SaaSHero acts as the outsourced inbound growth team for B2B companies, with one team owning strategy and execution across paid media, creative, landing pages, and reporting, all tied to CRM revenue data instead of form-fill counts.

The fair comparison problem disappears when one party is accountable for the entire chain from impression to CRM record. SaaSHero owns that chain end to end, including ad account, creative, landing page, conversion tracking, and CRM-connected reporting in HubSpot, Salesforce, or the client’s own CRM. One team owns the entire funnel, so the gap between the click and the pipeline record has a single accountable owner.

SaaS Hero: The client-friendly SaaS marketing agency that proves pipeline
SaaS Hero: The client-friendly SaaS marketing agency that proves pipeline

The retainer is flat and indexed to total monthly ad spend rather than channel count. Adding a channel, shifting budget from LinkedIn to Google, or testing Meta does not raise the fee. Channel-mix recommendations are made on evidence instead of pricing incentives.

SaaSHero has managed B2B SaaS paid media since 2018 and currently oversees roughly $16 million in annual ad spend across more than 100 clients. That scale makes the CRM-connected reporting credible because the benchmarks and patterns come from accounts in the same category, at similar spend levels, with comparable sales-cycle complexity.

SaaS Hero: Trusted by Over 100 B2B SaaS Companies to Scale
SaaS Hero: Trusted by Over 100 B2B SaaS Companies to Scale

For a VP of Marketing who must prove to a CFO that the agency is worth the fee, SaaSHero’s CRM-connected reporting answers the real question: which spend produced which pipeline, at what cost, against what payback period. That artifact survives a board meeting without requiring a rebuild from three conflicting sources.

TripMaster adds $504,758 in Net New ARR in One Year
TripMaster adds $504,758 in Net New ARR in One Year

For more on how the models compare at the B2B SaaS level, see B2B Marketing Agency vs In-House: 2026 Cost-Effectiveness.

See SaaSHero’s CRM-Connected Reporting In Action

Frequently Asked Questions

How Often Should You Run An Agency Versus In-House Evaluation?

Run the evaluation quarterly so it aligns with your board or sponsor review cycle and feeds directly into budget defense. Pipeline coverage, CAC payback, and LTV:CAC are quarterly metrics, not annual ones, so an annual review leaves you defending last year’s spend with last year’s data.

Quarterly reviews also give both models enough runway to show leading indicators such as SQL rate, landing page conversion rate, and pipeline trajectory before revenue results are available. That timing creates the only fair basis for evaluation inside a 90-day window on a 90-day sales cycle.

Does The In-House Versus Agency Scorecard Work For Smaller Teams?

The same scorecard structure works for smaller teams with a few adjustments. Weight controllable factors more heavily and use descriptive ratings instead of numerical scores when data volume is low.

A two-person in-house marketing team cannot be scored on the same pipeline coverage threshold as a five-person team because the denominator differs. An agency managing $15,000 per month in ad spend also should not be evaluated against benchmarks built from accounts running $100,000 per month, because smart bidding needs more data at higher spend levels.

The core structure remains the same: separate controllable from uncontrollable factors, score on CRM-connected outcomes, and use descriptive ratings. Calibrate thresholds and metric weights to the actual scope each party was hired to cover.

How Do You Handle Attribution Gaps And Long Sales Cycles?

Use multi-touch attribution, push lifecycle-stage events back to ad platforms, and separate primary from secondary conversions. Set attribution windows to match your actual sales cycle, such as 60–90 days for mid-market and 90–120 days for enterprise, instead of relying on the platform’s default 7-day or 30-day window.

The practical mechanism is an offline conversion import that fires CRM stage transitions such as MQL, SQL, Opportunity, and Closed Won back into the ad platform so the bidding algorithm learns from qualified outcomes rather than form fills. Without this loop, the platform optimizes for whoever fills out forms fastest, which is not the group that ultimately buys.

On the reporting side, run first-touch and multi-touch attribution models in parallel. Compare how channel credit shifts between them. The differences reveal which channels create top-of-funnel demand that last-click attribution systematically ignores and defunds.

What If The Agency And In-House Team Are Both Underperforming?

Consistent underperformance on both sides usually points to a structural problem. Start by auditing whether anyone owns the full chain from impression to CRM record.

If the agency owns the ad account but not the landing page, the in-house team owns the CRM but not the conversion tracking configuration, and a web contractor owns the landing page but not the campaign brief, then no single party is accountable for the outcome. In that setup, the evaluation will keep producing the same result regardless of who gets scored.

The fix comes from closing the ownership gap instead of replacing the agency or the in-house team. One party must be accountable for the whole path, or the scorecard will always find someone to blame and never surface the actual problem. SaaSHero is designed to resolve this structural gap by owning strategy, execution, creative, landing pages, and CRM-connected reporting as one team on one accountability line.

Read Next