Does Your Test Know What It’s Testing?

August 12, 2026      Kevin Schulman, Founder, DonorVoice and DVCanvass

Before you run a fundraising test do you specify what behavior is supposed to change?  Is it response rate, average gift, maybe both?

If you can’t answer that before the test goes out then that tells you something.

That “something” usually shows up in the test results that report on response rate, average gift, revenue, cost to raise a dollar and whatever other numbers happen to be hanging around.  A spreadsheet with enough columns can usually find you some good news.

So what’s the issue with column surfing for good news?  In one study of 389 randomized clinical trials, researchers found that trials that moved the goal post by changed at least one primary outcome between trial registration and publication reported treatment effects about 16% larger than trials that did not.  Allowing for what might feel like reasonable flexibility in what you analyze and report makes it more likely you’re manufacturing significant findings that aren’t there.

Suppose we’re testing something explicitly intended to make saying yes easier, perhaps we’ve reduced friction or changed the message. Our hypothesis is that more people will respond, so response rate is the primary outcome and average gift among responders doesn’t get an equal vote simply because it’s sitting in the campaign report.

Now imagine,

  • The control goes to 1,000 people. Twenty give $50, for a 2% response rate and a $50 average gift.
  • The treatment also goes to 1,000 people. The same twenty people give their same $50, but the treatment persuades ten additional people to give $20.
  • Response has increased from 2% to 3%. Revenue has risen from $1,000 to $1,200. Nobody gave a penny less because of the treatment.
  • Average gift, however, fell from $50 to $40.

If you report that the treatment “reduced average gift by 20%,” you’re describing the arithmetic accurately while describing the behavioral effect incorrectly. The treatment didn’t make donors give less. It created additional donors who happened to give smaller gifts.

This isn’t a statistical parlor trick or academic nuance. Once treatment changes response, you’re calculating average gift from two different populations of responders.  This is because whether someone gives at all is a very different set of motivational and behavioral conditions from how much a person gives.

So before the package drops, the email deploys or the ad starts running, write this sentence:

We believe [treatment] will change [response rate / gift amount / both] because [mechanism].

Then design the analysis around it.

If you’re trying to increase response, power the test for response rate and declare that the primary outcome. Average gift can be a secondary diagnostic or guardrail, but don’t let a random wobble overturn the result.

If you’re testing an ask string, anchor, upgrade mechanism or something else explicitly designed to change how much people give, gift amount belongs front and center.

And if you genuinely expect both to move, say so beforehand. The important part is deciding before you know which number looks prettiest.  A test is supposed to tell us whether our hypothesis was on the markt, not which metric won.

Kevin

Leave a Reply

Your email address will not be published. Required fields are marked *

Leave a Reply

Your email address will not be published. Required fields are marked *