What to Measure for Agency Case Study Verification Before Shortlisting a Digital Agency

Agency case studies are useful when they reduce uncertainty about a provider’s method, scope, and evidence. They are weak when a large number is detached from a baseline, a period, or a measurement path. Before shortlisting an agency, treat each case study as a set of claims to verify rather than as proof that the same outcome will occur for your business.

Define the decision you are trying to make

Write the capability that matters: qualified demand, paid-search efficiency, CRM repair, website conversion, sales handoff, or another defined job. Then describe the situation in which the agency would have to perform. An industry logo alone is not a relevance test; a different sales cycle, market, margin, or data environment can change the work materially.

Create a short comparison cohort. For every case study, record client type, problem, market, offer, channels, scope, team, and period. If the agency publishes only a broad vertical label, mark the missing detail instead of assuming similarity.

Score the evidence, not the design

Use the Agency Case Study Verification Scorecard with a 0–2 scale for each dimension:

| Dimension | 0 | 1 | 2 | | — | — | — | — | | Relevance | no comparable job or audience | partial similarity | same decision and operating context | | Baseline | no starting definition | verbal starting point | dated population, metric, and source | | Intervention | outcome with no method | tool or channel list | decisions, sequence, and owners explained | | Measurement | vanity or unnamed number | platform report only | traceable path to defined business outcome | | Time window | no period | broad or unclear period | explicit before, after, and observation window | | Attribution | causal claim without test | qualified correlation | design or reconciliation supports the claim | | Constraints | no limitations | generic disclaimer | specific exclusions, dependencies, and trade-offs | | Confirmation | no client or source trail | vendor-selected quote | permission, reference, or auditable source available | | Repeatability | one dramatic example | similar stories without method | pattern explained with conditions and failure cases |

The score is a screening aid, not a scientific benchmark. A case with a lower score may still be relevant, but the missing evidence becomes a question for the next conversation.

Check the claim language

List every objective statement: percentage change, revenue figure, ranking, cost reduction, speed improvement, or superlative. Ask whether the statement is express or implied, what a reasonable buyer would infer, and what evidence existed before publication.

The FTC advertising and marketing basics says advertisers need a reasonable basis for claims before they run and that evidence must support the claim as conveyed. This does not turn an agency review into legal advice, but it is a strong practical rule: request the definition, source, and calculation for a material result.

If the page uses a named client, logo, quote, or personal experience, review the FTC Endorsement Guides. Ask whether the person actually had the described experience, whether a material connection is disclosed where relevant, and whether the wording implies typical results that the evidence does not establish.

Trace the measurement path

Ask the agency to show the path from intervention to reported outcome. Depending on the service, that can include:

  • a campaign, content, or website change log;
  • analytics event and conversion definitions;
  • CRM lead, opportunity, and stage records;
  • finance or payment reconciliation;
  • exclusions, refunds, cancellations, and offline outcomes;
  • the attribution model and lookback window;
  • a list of concurrent changes and known confounders.

Google’s people-first content guidance asks for original analysis, clear sourcing, and demonstrable expertise. Use the same question with a case study: what did the agency actually learn or measure that a generic result summary does not show?

Do not confuse a platform-reported conversion with incremental revenue. Ask which records were accepted by sales, which progressed, and which were later rejected. If the agency cannot disclose confidential numbers, it can still provide definitions, ranges, a redacted ledger, or a client reference. “Confidential” should narrow the claim, not make every claim unverifiable.

Test the alternative explanations

For each material result, write at least one credible alternative explanation. A conversion-rate change may follow a pricing change, a different audience, a season, inventory availability, a tracking repair, or a sales-capacity change. A ranking change may reflect a search-system update or competitors leaving the result set. A lead increase may include duplicates or low-intent submissions.

Ask how the team separated those explanations. The answer may be a controlled test, a cohort comparison, a reconciliation ledger, a clearly bounded observation, or an admission that causality is unknown. An honest unknown is a stronger signal than a confident claim without a method.

Request targeted verification

Send the agency five specific questions for each shortlisted case:

  1. What was the starting condition, and how was it defined?
  2. What changed, in what order, and who owned each change?
  3. Which source systems support the reported result?
  4. What did the case not measure or prove?
  5. May we speak with a client or review a redacted evidence record?

Compare the answers with the public page. Differences are not automatically disqualifying, but they should be explained. If the public story says “revenue growth” and the evidence only supports “platform conversions,” the claim should be narrowed before the agency is considered for a revenue-critical engagement.

Use the scorecard to choose the next action

Classify each case as verified enough for shortlist, useful but needs clarification, not comparable, or hold. Record the reviewer, date, evidence requested, response, and unresolved risk. Do not average away a fatal gap in identity, permission, data integrity, or scope fit.

Use the shortlist to shape the discovery call. Ask how the agency would reproduce the measurement path in your environment, what access it needs, which decisions remain with you, and what it would stop if the evidence does not support expansion. Case studies should inform a testable engagement, not replace one.

Final pre-shortlist gate

Shortlist an agency only when at least one relevant case explains the problem, intervention, evidence, and limits clearly enough for your team to ask an informed follow-up. Keep missing baselines, unverified claims, and vendor-reported outcomes visible in the decision record.

The best provider is not the one with the most dramatic portfolio page. It is the one whose evidence can be inspected, whose assumptions can be discussed, and whose proposed measurement contract survives contact with your CRM, sales process, and financial reality.

Your reaction

How did this article land?

Choose one reaction. You can change it anytime.

Email verification required

Write for Scale Orbit

Turn practical experience into a public body of work

Share useful lessons about revenue, marketing, analytics, CRM, conversion, and growth. Build a visible author profile and learn what resonates with practitioners.

  • Public author profile and publication archive
  • Editorial support for your first article
  • Views, reactions, followers, and topic discovery
  • Free publishing with clear moderation rules

Email verification is required. Every first article is reviewed. Publication, rankings, traffic, leads, and revenue are not guaranteed.

Write

Discover more from Scale Orbit | Full-Service Marketing Management

Subscribe now to keep reading and get access to the full archive.

Continue reading