Day 41 · Phase G

"Make the list longer" is the one effort in prospecting that reliably does nothing

Cold start · ideal customer profile · choosing channels · the numbers game·4 moves · 4 diagrams
What's scarce at a cold start was never names. It's a sentence you can say that is true of this person and no one else. And the numbers game is a game about variance — it rewards people willing to change audiences and punishes people who polish wording.
Prospecting is the work of narrowing "everyone alive" down to the few dozen people worth an hour of your time, and getting them to start a first conversation. Four moves take it apart, in order. One, subtract your profile from wins minus losses — a profile that can't shorten your list isn't a profile. Two, run the first thirty conversations for their words, not for a sale. Three, compare channels by minutes per real conversation, and run only two at a time. Four, treat the numbers game as a variance game — small samples can't detect your optimizations, so bet on step changes and track absolute counts.
MOVE 01

A profile is only worth writing if it excludes people The test of a customer profile is how many names it deletes.

discriminationpositive testingscorecards
"Mid-sized, growing fast, cares about efficiency, has budget." The problem with an ICP (ideal customer profile) like that isn't that it's wrong — it's that it excludes nobody. A trait is worth something only if the share of your won deals that have it differs sharply from the share of your lost deals that have it; a trait running equally high on both sides is self-consolation in list form. So a profile can't be reasoned outward from your product. It has to be subtracted from a side-by-side comparison of the last twelve months of wins and dead deals.
Same set of deals, four candidate traits share of won deals share of lost deals New owner in last 6 mo. 89% 10% gap 79 → keep Already running the numbers 78% 15% gap 63 → keep 10–200 employees 90% 85% gap 5 → cut Has an online store 100% 95% gap 5 → cut The bottom two aren't wrong. They're empty: they describe your wins and your losses equally well. Acceptance test: adding this line, does my list of 3,000 shrink? Still 3,000 → delete the line. (Numbers illustrative — replace them with the proportions you counted in your own two columns.)
Setup: you're about to run a round of outreach to "small and mid-sized e-commerce companies."
✗ Filtering on criteria that exclude nobody

The criteria are "has an online store, 10–200 employees," and the list comes back with three thousand names. Neither criterion rules anyone out, so all three thousand emails have to say the same thing — a sentence true of everybody, and therefore true of nobody.

✓ Putting last year's wins beside last year's dead deals

Nine wins, twenty dead. Only two traits differ sharply across the two columns: eight of the nine wins had changed the owner of this area within six months, against two of the twenty losses; seven of the nine were already running the numbers on it in a spreadsheet, against three of the twenty. The list drops from three thousand to sixty — and every email can finally name something true of that one company.

Read-aloud template · subtracting a profile in three steps

1. Two columns. Last twelve months of wins in one, dead deals in the other (losses and no-decisions both count as dead). Small sample — go back twenty-four months.

2. Count both proportions per trait. What share of wins has it, what share of losses. Gap under about twenty points → delete it on the spot. It has no power to discriminate.

3. Write what survives as a 0/1 scorecard. One point per trait, score the list, work only the top band. Don't substitute "this one feels like a fit" for the score.

Why it works

A profile is a classifier, and a trait carries information only when its rate differs between the two classes; equally high on both sides means zero discriminating power. Meanwhile people write profiles by recalling the customers who bought — sampling positive cases and never the negative ones is a default cognitive strategy that has been measured over and over, and it grows exactly the kind of profile that is wide, comfortable, and unable to separate anything.

  • Hard evidence · people only generate examples that fit their hypothesis: in Wason's (1960, Quarterly Journal of Experimental Psychology) 2-4-6 task, subjects guessing a number-sequence rule almost exclusively proposed sequences consistent with their current guess and rarely tried one that could overturn it, leaving them stuck for long stretches on a rule that was too narrow and wrong. Classic and repeatedly replicated. Your ICP is that guess, and your lost deals are the test you never ran.
  • Hard evidence · the positive test strategy: Klayman & Ha (1987, Psychological Review) made that precise — people default to examining cases where the hypothesis holds. Under some structures it's an efficient heuristic, but when the task is telling two classes apart it fails systematically. Prospecting is the second kind of task.
  • Hard evidence (meta-analysis) · the scorecard beats the gut: Grove et al. (2000, Psychological Assessment) pooled 136 studies comparing mechanical prediction (formulas, scorecards) against expert clinical judgement. Mechanical was on average equal or superior, and clearly worse only in rare cases. So the ugly 0/1 table will probably outperform a veteran rep's "I have a good feeling about this one."

Boundary: with fewer than five wins, everything this subtraction produces is noise. Skip the proportions and use one qualitative signal instead — what had just happened inside the company the last time someone bought. Profiles also expire: recompute when the product or the market moves, rather than running a table you subtracted two years ago. (Whether one specific name on the list deserves more of your time is the qualification piece.)

MOVE 02

The first thirty conversations exist to harvest their words Not to close anything. To bring their sentences back.

processing fluencyphrase poolcold start
At a cold start what you're short of isn't names — it's that you can't say a single concrete thing about these people. Sending three hundred emails at that point just mass-prints a hypothesis you haven't found the language for. Do thirty real conversations first (twenty minutes is plenty), with exactly one job: copy down the words they use when they describe the problem. When one phrase has come up six times in thirty people, you have your opening line — and you didn't write it, they did.
Four steps that refine an opening line out of their mouths 30 conversations, 20 minutes each not selling copy their nouns and verbs verbatim never translate phrases heard 3+ times go in a pool ranked by count top phrase = your first line not one word changed ✗ Your wording (you wrote it) "We provide end-to-end inventory optimization that lifts turnover." ✓ Their words (you copied them) "Does month-end stocktake still shut you down for a day?" Every word on the left is correct, and not one of them is his. Test: can your opening line be found verbatim in your notes? If not, you are still talking to yourself. Thirty is enough — by the low twenties, new phrases have essentially stopped appearing.
Read-aloud template · booking and using the thirty

1. The ask: "I'm working on something around X, I'm not selling anything, I'd like to hear how you handle it today — twenty minutes." And then genuinely don't sell. Sell once and that door shuts permanently.

2. Ask about the past, not the future: "When did you last run into this? What did you actually do about it?" (The discipline for not getting misled by polite answers is the Mom Test material.)

3. Write it down verbatim: nouns and verbs exactly as spoken. Translating into your own vocabulary destroys the entire value of the conversation.

4. Count: any phrase heard three or more times goes in the pool, ranked by frequency, and the top one becomes your opening line unedited.

Why it works

Using words the other person already uses reads as "this one gets it" more reliably than your more accurate phrasing does, because familiar words take less work to process — and that feeling of ease gets misread as true, credible, warm. This is processing fluency: people cannot tell "I understood this easily" apart from "this is right," so they treat the first as evidence of the second.

  • Hard evidence · easier to process means better liked: Reber, Winkielman & Schwarz (1998, Psychological Science) raised only the processing fluency of a stimulus via contrast and priming, leaving content untouched, and liking rose significantly — with subjects entirely unaware of the cause. Peer-reviewed, solid.
  • Hard evidence · fluency even shifts judgements of truth: McGlone & Tofighbakhsh (2000, Psychological Science) found that semantically equivalent aphorisms were rated as more accurate descriptions of human behaviour when they rhymed. Fluency doesn't just move liking; it gets counted as evidence that the statement is correct.
  • Hard evidence (meta-analysis) · mere exposure: Bornstein (1989, Psychological Bulletin) pooled more than 200 studies confirming the exposure effect — familiarity alone produces liking, and it holds even without conscious recognition. Translated to outreach: in the first second a stranger isn't judging your logic, they're judging whether the sentence parses.

Boundary: these experiments measured liking and truth judgements, not reply rates, so the crossing into sales deserves a discount — they give you a direction, not an effect size. The reverse risk is real too: fluent isn't correct, and their own phrasing may be repeating a wrong self-diagnosis. So use their words in the opening, where the job is getting them to talk, and never in the diagnosis, which only follow-up questions can do.

MOVE 03

Compare channels by minutes per real conversation Reach is large everywhere, so it can't be the thing you choose on.

explore–exploitunit costonly two at a time
Most people pick a channel by how many people it reaches, but that number is large on every channel, so it can't function as a choice at all. The unit that decides it is your minutes divided by the real conversations you get back — a channel producing three twenty-minute conversations beats one reaching thirty thousand people. And don't run five at once: each gets too thin to teach you anything, and you'll never have the nerve to cut any of them.
Minutes this channel costs you per twenty-minute conversation Referral from a client 18 min Where you publish 45 min Bulk cold email 120 min Pure cold calling 200 min (Illustrative — replace with two weeks of your own log. Reach never enters this chart.) Four candidate channels, how to spend four weeks Weeks 1–2 equal small bets A · 25 each B · 25 each C · 25 each D · 25 each Weeks 3–4 cut the worst two winner A · double the volume winner C · double the volume One number decides the cut: real conversations from this channel. Opens, likes and new followers don't count.
Read-aloud template · the four-week channel trial

1. Pick four candidates and give each an identical small budget (25 touches each, or two hours each). Don't start by going all in on your favourite.

2. Score one number only: real conversations this channel produced. Opens, likes and followers stay off the sheet — none of them convert into revenue.

3. Cut two in week three and push all the time into the survivors. Retest a cut channel with a small bet six months later.

4. Note one more thing: can I be recognised here? Channels where you have work, shared people or an identity get cheaper over time; a purely cold channel's unit cost is a constant.

Why it works

This is a standard explore–exploit problem: each candidate channel is a slot machine whose true payout you don't know, and you're allocating between trying more of them and concentrating on the best one so far. The optimal shape is always the same — equal small bets first, then fast concentration — which is neither going all in at the start nor spreading five ways forever. There's a second layer: channels where you have an identity have decreasing unit cost, because what you leave there gets seen repeatedly, while a purely cold channel starts from zero on every single touch.

  • Mathematical result (not psychological evidence) · multi-armed bandits: Gittins (1979, JRSS Series B) proved the optimal index policy for the discounted multi-armed bandit. That's a theorem, not an empirical regularity, and the practical "spread small, then concentrate" recipe is a crude approximation of that family of results.
  • Industry practice (not experimental, weaker evidence): Ross & Tyler's Predictable Revenue (2011) argues for splitting lead generation and closing between different people, on the grounds that the two require different rhythms and skills. Widely adopted — but what supports it is company case studies rather than controlled trials, so hold it as a hypothesis, not a conclusion.
  • Moderate strength · saturate one small circle before expanding: the recurring pattern in Rogers's Diffusion of Innovations survey is that new practices work through a small, homogeneous circle before spreading outward. A synthesis of observational research — direction credible, force discounted.

Boundary: channel economics drift, and a channel that is cheap this year may be crowded next year, so retest what you cut with a small bet every six months. And "real conversations" fails as a counting unit in very low-price, very high-volume businesses — nobody needs a twenty-minute call before buying a coffee. There, count conversion steps instead.

MOVE 04

The numbers game is a variance game, not a diligence game At low reply rates, differences inside a small sample are noise.

law of small numbersstatistical powerstep changes only
"Outreach is a numbers game" is true, and nearly everyone reads it backwards. It doesn't mean send more. It means at a low reply rate, differences inside a small sample are noise. Fifty emails yielding one reply against fifty yielding three looks like a threefold win; the two intervals sit almost entirely on top of each other. So: stop polishing wording and bet on step changes — a different audience, a different concrete fact about the person in your opening line, a different channel. Then track absolute counts only: how many real conversations this week.
One outreach round, 95% intervals for two arms of 50 the whole stretch where they overlap Arm A: 50 sent → 1 reply (point estimate 2%) 0.4% 10.5% Arm B: 50 sent → 3 replies (point estimate 6%) 2.1% 16.2% 0% 5% 10% 15% 20% Wilson intervals. B's point estimate is triple A's, yet the bands lie almost on top of each other. Telling 2% from 4% reliably takes upwards of a thousand sends per arm. Your sample detects nothing small. Consequence: small changes are untestable, so only make big ones — change the audience, not the adjectives.
Read-aloud template · four rules for the numbers game

1. Fix the batch: one audience, one sentence, minimum thirty, and no edits until it's finished. A batch you changed halfway through is void.

2. Change one variable at a time, and only big ones: audience > the concrete fact about them in your opening > channel > wording. Wording ranks last because its effect is smallest and its sample demand highest.

3. Put only absolute counts on the wall: sent this week, real conversations this week. The second number is the one worth displaying; a reply rate's decimal place is not.

4. Commit to a volume before you're allowed to judge — say a hundred attempts, no conclusions and no changes of mind before then. Otherwise what you kill is the unlucky batch, not the bad one.

Why it works

The lower the reply rate, the less information a given sample size can carry — one reply's worth of random fluctuation can double the ratio. Yet people systematically believe small samples represent the population, so they read "this opening doesn't work" off twenty emails and kill an approach that was actually fine. There's a second layer, and it's emotional: every rejection genuinely stings, so people substitute "let me tweak the wording" for "let me send another batch" — because tweaking wording requires facing no new rejection at all.

  • Hard evidence · the law of small numbers: Tversky & Kahneman (1971, Psychological Bulletin) showed that even statistically trained researchers overestimate how representative small samples are, and consequently conclude too early and abandon effective approaches. Peer-reviewed and classic. Staring at the reply rate on twenty emails is exactly this.
  • A calculation, not a citation · how big a sample: by the standard two-proportion power calculation (α = 0.05, power = 0.8), reliably distinguishing a 2% from a 4% reply rate needs upwards of a thousand sends per arm. That's computed rather than cited — but its implication is hard: your sample size cannot detect any small change.
  • Contested — treat as a mechanism hint only: Eisenberger, Lieberman & Williams (2003, Science) found in the Cyberball paradigm that social exclusion raised activity in dorsal anterior cingulate cortex (dACC) and anterior insula, regions overlapping with the distress component of physical pain — the sting of rejection isn't purely metaphorical. But the strong version has been weakened: Woo et al. (2014, Nature Communications) used multivoxel pattern analysis to show that physical and social pain are separably represented within those same regions. Fine as an explanation for why outreach is hard to sustain; not a settled fact.

Boundary: all of this holds only for outreach with low reply rates, modest deal sizes, and enough volume to accumulate a sample. If your total addressable set is forty companies, the statistics collapse entirely — the right move there is to research and design for each one separately, treating outreach as forty projects rather than one batch.

Your Day 41 Action

One hour to trade "a very long list" for "a very short list where you can say something about every name."

One (20 minutes) · subtract the list: write out the last twelve months of wins and dead deals, count both proportions per trait, find the two traits that run high only among the wins, turn them into a 0/1 scorecard, score your current list and keep only the top band.

Two (20 minutes) · book three conversations: pick three names from that band and send "I'm not selling anything, I'd like to hear how you handle this today — twenty minutes." Sending counts as done — whether they say yes isn't your score.

Three (20 minutes) · fix the batch: write down the one number you'll track for four weeks (real conversations this week) and one commitment: no conclusions and no changes of mind before a hundred attempts.

Boundary: this machinery is for businesses with hundreds or thousands of possible customers. With a few dozen, skip the batches and the ratios entirely — research each one separately instead. One more thing: the most valuable of your thirty conversations are usually the people who say plainly "we don't need this." What they tell you is where the profile should be subtracted next.
Think It Through
1. I'm an individual — freelancing, or job hunting. I have no "wins and losses" to compare. How does this apply?
You have them, you just never wrote them down. Wins are the people who actually paid, made an offer, or talked seriously to the end; losses are the people you talked to who went quiet. Three or four of each is enough to subtract one line.

Put them side by side and look for what appears only among the ones that worked. The category that shows up most often is a trigger event: they'd just changed who owned this, they'd just had something go wrong, or there was a hard date pressing on them. Triggers beat industry and company size, because they don't determine whether you're a fit — they determine whether it's urgent now, and urgency is why the door opens at all.

At that sample size don't compute proportions. Write one sentence instead ("look for people who've just inherited this area and haven't built a process yet") and test it against your next ten. As for the thirty conversations, individuals actually book them more easily: without a company badge, people's guard is far lower.
2. Isn't cold email or cold calling just spam? I can't get comfortable with it.
The test isn't "they didn't invite you" — by that standard every act of commercial initiative is spam. It's these three:

One, can you say something true of this person and nobody else? If you can't, it is a blast, and it is spam. Two, how expensive is it for them to refuse? An email deleted in two seconds and a phone call at ten at night are not the same object. Three, would you be comfortable letting them see exactly how you reasoned about contacting them? If not, some part of you already knows the answer — this is the sharpest of the three.

Clear all three and cold outreach is ordinary commercial behaviour, no guilt required. One practical addition: anyone who says plainly "not interested" comes off the list and genuinely never hears from you again. The value at stake was never this one deal — it's whether you can still be recommended inside that circle, and that pool is worth far more than this round's list.
3. I've sent a hundred following all of this and got zero replies. Is the list wrong or the message wrong?
Find out which layer is empty first. A hundred sends and zero replies is almost never a wording problem — wording moves you between 2% and 4%, not between 0 and 2%. Zero usually means one of exactly three things: wrong audience (these people simply don't have the problem you're describing), wrong trigger (they have it, but nothing makes it urgent right now), or wrong channel (the message was never seen — it went to spam, or nobody reads that account).

Check them in reverse order, cheapest first. Ask one or two people you actually know: "did my email reach you, did you open it?" Ten minutes rules out a whole class of cause. Then check the trigger: does your list have a "what recently happened" column at all? Without one you're pitching people who aren't in a hurry. Change the audience after that, and touch the wording last — always last.

And if you've swapped audiences twice and it's still zero, you may be selling a problem nobody is in a rush to solve. What that calls for is going back to the thirty conversations, not attempt one hundred and one.