Lead Scoring and Outreach at 12,000 Leads a Week, With Reply Rates Up 42%
Enrichment and scoring are the easy half. The reply rate moved because the messages stopped sounding like they were generated, which took more work than the model did.
- Client
- B2B revenue team
- Industry
- Sales technology
12,000+
Leads processed weekly
+42%
Reply rate
Automated
Discovery to CRM sync
Reusable
Bot framework across teams
The problem
The team's pipeline work was three jobs pretending to be one: finding prospects, working out which were worth contacting, and writing something worth reading. Each was done partly manually, and the third was done badly under time pressure — which meant good leads received generic messages and converted at generic rates.
Volume made it worse rather than better. Scaling outreach with templates raises send count and lowers reply rate, and past a point the two cancel out.
What we built
A discovery and enrichment pipeline that sources prospects and appends firmographic and behavioural data from multiple providers, deduplicating and reconciling conflicts rather than trusting whichever source answered last.
A predictive scoring model trained on the team's own closed-won history, so ranking reflects what converts for this business rather than a generic ideal-customer template. Segmentation runs off the same features, which keeps the scoring and the messaging strategy consistent.
A generative outreach layer that composes each message from the enriched record — the prospect's stated problem, their stack, their segment, their trigger event — rather than filling three variables into a template.
Two-way CRM synchronisation so scores, segments and every touch write back automatically, and an internal builder framework so other teams can deploy their own outreach bots on the same rails without a new engineering project each time.
The part that needed care
Personalisation at volume fails in a specific way: it reads as automated, and a message that reads as automated performs worse than an obvious template because it feels like a trick rather than a shortcut.
We spent more effort on the composition layer than on the scoring model — constraining structure, forcing specificity from the enrichment data, and rejecting any generated message that could have been sent to a different company without editing. Anything that failed that test did not go out.
Results
The platform processes more than 12,000 leads a week through discovery, enrichment, segmentation and scoring, with outreach generated per prospect rather than per segment.
Reply rates improved 42% against the previous approach at comparable volume. The pipeline did not get bigger; it stopped wasting its best leads on its worst messages.
The builder framework turned out to matter as much as the results. Other teams adopted the same pipeline for their own campaigns without engineering time, which is what made the improvement organisation-wide rather than one team's advantage.
What we'd flag
Scoring trained on your own history needs enough history to learn from. A team with a few dozen closed deals should use rules and revisit the model later; fitting one to sparse data produces confident nonsense.
We also capped sending volume deliberately. The temptation once this works is to increase throughput until reply rates collapse, and holding volume steady while improving relevance produced a better result than any volume increase we modelled.









