GTM VaultPro

Library/Show Me Your Stack 4

From Win List to Win-Loss Gap: Building the Propensity Layer

Why Jai Toor stopped running signal work on wins alone and built a Claude Code workflow on the win-loss gap, compressing a month and $2,000 into 30 minutes and $10

Jai Toor2026-04-268 min readWatch on YouTubeSubstack post

This episode: Jai Toor walks through the Niche Signal Discovery skill built on top of Claude Code and DeepLine, a workflow that pulls wins and losses from HubSpot, mines the differential signal between them, and exports a verified-lead outbound campaign into Lemlist end-to-end from a single terminal. The system compresses a month of manual propensity work and nearly two thousand dollars of enrichment spend into thirty minutes and under ten dollars. The breakdown below maps the build back to the Revenue Architecture.

The Signal Commoditization Gap

Most GTM teams run their signal layer on the same three inputs: hiring posts, funding rounds, and tech stack changes. Everyone scrapes the same sources. Everyone ingests the same data into the same enrichment pipelines. Everyone reaches out on the same week. The signal stops being a signal the moment everyone is pulling from the same place.

This is not a data gap. It is a sourcing gap. The data sits on every target’s public surface. What is missing is a method for finding which specific patterns on that surface correlate with wins, and the compute to test that against every account at a cost that makes the work economical.

Jai has spent the last year building the action layer that closes that gap. Before DeepLine, Jai spent five years on growth at Uber, then built data and product at Capchase, Datafold, and VeriShop. He co-founded DeepLine, an integration layer that sits underneath the agent and handles every API call in a signal workflow through pay-as-you-go credits. He sits at the intersection of growth engineering and data infrastructure, which is exactly where this problem lives.

In this episode, Jai walks through the full build. First, the Niche Signal Discovery skill running inside Claude Code. Then the infrastructure layer that makes the skill possible: DeepLine’s connectors to HubSpot, Lemlist, and every data source the workflow needs to reach. The argument for why this is architecture and not a tool is that the same shape runs at every layer of the signal stack once the context layer exists underneath.

The Win-Loss Foundation: Why Most Signal Work Starts in the Wrong Place

Jai’s first structural point is that most signal discovery runs on the win list. Take closed-won accounts, find the common keywords on their websites, target more accounts that look like them. The output is generic ICP work dressed up as signal discovery. Everyone selling to consumer fintech will find “fintech” on the website. That is a description, not a signal.

Jai inverts that. The workflow pulls wins and losses from HubSpot together, not just wins. Both cohorts are accounts the team engaged. Both cohorts were qualified enough to reach opportunity stage. The difference between them is the real selection criterion the market is using, whether the team designed it or not.

In the cybersecurity case walked through on the episode, the dataset was seventy-four deals. Wins and losses combined. Claude Code scraped the customer websites of every account on that list, generated keyword hypotheses programmatically, tested them against the scraped content, and surfaced the keywords that separated wins from losses. The positive signal for that ICP was FDIC mentioned on the company website. Specific. Testable. Almost certainly a consumer fintech company with regulated retail deposits.

The previous version of this model had used “open an account” as the proxy for the same concept. Directionally correct. Categorically worse. The specificity of the surfaced keyword is the difference between a signal that compounds and one that decays into a category filter.

Accounts that differentiate cleanly between wins and losses are the only accounts worth running downstream signal monitoring on. Everything else produces noise in the outbound queue.

SCREENSHOT: The Niche Signal Discovery skill starting up in Claude Code. 74 deals identified in HubSpot. Wins and losses pulled together as the foundation for the differential analysis.

The Differential Signal Layer: Positive and Negative Fit

The other half of the architecture is the negative signal layer. Most scoring systems only reward positive fit. Keywords that correlate with wins get a positive score and everything else gets ignored. Jai’s system builds both sides.

Two examples from the episode. SOC 2 Type 2 compliance on the target’s website was a positive for B2B buyers. Intuitive. The counterintuitive one was advanced engineering culture with heavy in-house tooling. Engineering-dense companies often look like good prospects on paper because they have the technical maturity to evaluate the product. The model surfaced the pattern that they close less frequently because they tend to build rather than buy.

The consequence of running only positive fit is that the BDR team spends its time on accounts that look like wins but behave like losses. One of the enterprise customers in Jai’s data had been spending 44% of their opportunity time on bad-fit accounts. Not because the BDRs were unskilled. Because the scoring model had no way to surface the structural signals that kept those accounts from closing.

Fix the negative signal layer, reclaim the 44%, and close rate moves 15 to 17 percentage points before anything else changes.

Figure 1: Win-Loss Differential Scoring Architecture. Wins and losses are pulled from HubSpot as a single cohort and run through the same pipeline. The composite score is the gap between the positive and negative fit layers, not the positive layer alone, which is where most scoring systems stop.

The Niche Signal Discovery Skill: One Command, End-to-End

With the scoring architecture in place, the actual skill runs from the terminal. One command. End-to-end execution.

The agent pulls every won and lost deal from HubSpot through DeepLine. It scrapes the customer websites in parallel. It generates keyword candidates from the scraped content, tests each candidate against the win and loss pools, and scores the ones that differentiate. It builds the composite propensity model. It runs the model against a hundred unqualified target accounts pulled from the enrichment source, scores them, and drops the top tier into a Lemlist campaign with verified emails and generated first-line personalization.

The output is a markdown doc. Five takeaways. A signal model with positive fits and negative fits, each scored. A hundred verified leads exported into Lemlist, with sequence copy generated, sender deliberately unattached so the operator has to confirm before the send.

Human in the loop at the send step is intentional. Every agent in Jai’s stack gets a confirmation gate before any external action fires. The Lemlist sender is unattached by default. The email drafts wait for approval. The CRM sync confirms before writing. This is the trust-building layer that gets skipped most often, and the one that determines whether a team keeps using the system after the first misfire or abandons it.

SCREENSHOT: Anti-fit signal table from the markdown output doc. Each negative signal scored by lift (under 1.0x = appears more in losses than wins) with risk indicator and structural reason for why the pattern correlates with closed-lost.

The Compression: Why Bespoke Enrichment Beats Off-the-Shelf

Jai ran this exact workflow manually for the same cybersecurity customer about six months before Claude Code and DeepLine made the end-to-end version possible. The comparison is the case for the architectural shift.

The manual version took over a month. It cost nearly $2,000 in enrichment spend. The team had to sit together and hypothesize roughly forty candidate keywords, then pay for AI to extract each one from every website in the dataset, then manually figure out which of the forty actually mattered. Most did not.

The current version runs in under thirty minutes and under $10. The keyword hypothesis step, which was the unscalable manual bottleneck, is now generated by Claude Code based on the content the agent reads. The extraction runs through DeepLine. The scoring runs in the same session. The outbound campaign writes itself into Lemlist at the end.

$2,000 to $10. One month to thirty minutes. Same outcome. Different infrastructure.

The output is not marginally better than the manual version. It is more specific. The manual version surfaced “open an account” as the proxy for consumer fintech. The current version surfaces FDIC. That specificity compounds at every downstream step: the outbound filter is tighter, the copy is more pointed, the reply rates move. Reply rates on campaigns running on Jai’s signals land at roughly 2x the baseline. A CPG customer who flipped from scraping retail websites to scoring every company with a store locator saw the same 2x lift on cold email and LinkedIn. An enterprise customer saw 15 to 17 point higher close rates once the negative-signal layer was active.

The enabling variable is Claude Code plus DeepLine. Jai’s not writing production infrastructure code to build this. DeepLine handles the API integrations through pay-as-you-go credits. Claude Code handles the orchestration. The bottleneck moves from engineering execution to problem specification. What are we optimizing for, and what does “won” actually mean in the context of this business.

The Adoption Curve: Who Is Actually Building This

The BDR function is not disappearing. It is being redeployed.

Two patterns show up in Jai’s customer base. The first is larger BDR teams with higher per-rep ROI. A BDR who used to touch 20 accounts a day is now touching 40 to 80. The raw information gathering that used to fill the calendar is gone. The rep spends the reclaimed time on calls, campaigns, and multi-threading. Revenue per rep goes up and total rep count also goes up because the unit economics of hiring another BDR improved.

The second pattern is BDR removal. One PLG customer stopped using the function entirely. AEs close the qualified PLG leads plus the automated outbound flow directly. No BDR handoff in the middle. This works when the ACV is low enough that the AE motion is economical and the close is AE-friendly. It does not work in long-sales-cycle enterprise environments where the BDR still does orchestration, phone work, and multi-threading that AEs are not positioned to run.

Cold calls still work. Jai’s customers in restaurants, CPG, and data analytics are getting meaningful lift from phone outbound, including to titles that were supposed to have moved online years ago. Chief data officers and VPs of data still pick up the phone when the signal is specific enough to justify the call. The constraint is not the channel. It is whether the signal upstream of the channel is differentiated.

Full transcript

Machine-generated transcript from the episode video. Speaker labels are not included and some names and product terms may be transcribed phonetically.

[0:00] Welcome to show me your stack. Today I'm with Jai Tour, co-founder of Deepline. Before Deepline, Jai spent five years on growth at Uber, then built data and product at Capchase, Dataf Fold in Vera. He sits where growth engineering meets data infrastructure, which is exactly where today's problem lives. Most GTM teams run signal layers on the same three inputs. Hiring posts, funding rounds, tech stack changes. Everyone scrapes the same sources. Everyone reaches out on the same week and the signal stops being a signal. Jai is going to walk us through how he uses cloud code and deepline to find the signals nobody else is running. Jai, let's see it. Awesome. Thanks for having me, Rick. As I'm going to share my screen here and uh yeah, so very tactically I run every single thing that I do through cloud code. I send emails through cloud code, email follow-ups. I do not leave this terminal interface for anything. Um, so what I'm going to run here is what we call our niche signal discovery skill.

[1:01] And I can make this a little bit bigger. And so I can uh describe it first and then show you. So niche signal discovery will go through the wins and losses. Go through our HubSpot wins and losses and then generate a uh summary doc and outbound campaign with 100 verified leads and one list. And so that's going to start churning. And what it's doing is we have connectors in the background through our tool deepline that connect to HubSpot that connect to Lemlist and then find wins and losses. So it has 74 deals in ours. This is going to run on deepline but I can show you the output for a cyber security company that we worked with and uh cool. So what the actual output comes out as is uh we typically share it with customers as a markdown doc just because that's easily sharable. I can run this on a company before actually speaking with them. And so it tries to come up with the five like really really key takeaways. What's the number one signal to look for? What

[2:03] does their ICP actually look like? What we found is what people describe their ICP as. And then what the data might show that their ICP is uh differs a little bit. And there's a lot of different reasons for that. Some are strategic like, oh, we don't want to go into healthcare anymore. Some are actually just things that they don't realize that they're optimizing for when they actually look at wins versus what they think their wins should be. So that's always an interesting part for customers. And then we'll give them actual contacts and leads with verified emails all in one prompt. So this is the output of what we did here. What Deepline actually adds to this system is it gives you all the API integrations that you need. You only sign up for Deepline. you get credits, pay as you go, and then it will use those credits to go find the right information from every data source. So, what it actually did is it scraped all of the wins and losses customers websites and find keywords that are unique and differentiated between wins and losses.

[2:58] So, it's really important that we also have losses in this. If you only do wins, you get pretty generic information. If you do wins and losses, those are people that you talked to or tried to have an opportunity with. And the difference between the two is where you you find a lot of alpha. So, what we've got here is keywords that were on customer websites that indicated that they were more likely to be a win versus a loss. So, this one like FDIC being mentioned, this customer's ICP is consumer fintech. So, that generally makes sense. That is keywords on the website. And then it might be who they were hiring before they closed. Uh, so if they were hiring fraud leadership or financial operations, those are good fits. Uh, and it breaks down all of these into really like tactical scoring models for you. and also has negative fits. So what was really common on losses that might indicate a company you shouldn't actually spend time on. Uh the top ones sock two type two this is a proxy for B2B. B2B companies generally prefer this. So that's something that they already kind of know but very tactical way to know if this is someone you should spend time on. And then this one was interesting. It's if they're

[4:00] building a lot of things in house and have advanced engineering culture that might not be the best fit for you guys as well. I'll stop there. Rick, any anything in particular stand out or not make sense? I mean, this looks great. If you can just um kind of tell us about traditionally how this how this workflow would have been run and how many people it would have taken would it have even been possible? Yeah. So, we did this exact workflow for this customer uh about 6 months ago and it took a little over a month and the process was flipped around, right? So, we took their entire total addressable market, went and found signals that we thought would be relevant, and then frontloaded all of it. So, we had to spend way more money to actually get all the signals. So, I think we spent over $1,500 um on just like raw signal scraping for that particular task. And then we processed it and found not quite we didn't find everything in here because the keyword discovery piece coming up with keywords and figuring out

[5:01] what is actually useful on a website that piece we did manually before. Now the cloud code system is actually generating a lot of these and then figuring out if they're useful. So it took us over a month and took almost $2,000 in terms of enrichment costs and now we do it in less than $10 and in about less than 30 minutes and we get to a very similar outcome. So the outcome that you get is this almost like an outbound score of what uh keywords are mentioned on the website tech stack what they're hiring for what might be a negative signal and then you kind of pass this through your organizational like expertise and say is this actually what we want to optimize for in the initial score that the the system delivered it said HIPPA compliance was a positive uh factor but the company had already made a decision to kind of depprioritize healthcare as a a vertical for unrelated like regulatory reasons.

[5:53] So that was something the model might say is true, but you actually decide is not not relevant. Yeah. So I would say it saved about two months overall and the quality is better and we use this pretty consistently. When we run it on a customer and just do this scoring for their outbound, we've averaged about 2x increase in response rates or open rates. a little bit that can be attributed to just like better copy uh being generated by the system, but overall really significant outputs, really low time investment, a much lower cost investment. Can you expand on the cost investment considering this probably wouldn't be possible without technology, right? Yeah. Yeah.

[6:34] Not to mention um the headcount reduction. Absolutely. Yeah. I I don't think the companies that we worked with could have done this without an expert previously. You needed to kind of know how to build a propensity model, know how to turn unstructured data into structured data and tie it to an account. So like very simply like you make a big table and say if this website has document authentication as a keyword, give it a one versus a zero. And you have to do that for thousands of different signals. That's kind of what cloud code is doing in the background for you automatically. and that data acquisition piece because we didn't know what signals mattered we the best solution that you can come up with is sit with the team hypothesize on like what they think is important for this particular customer we got a list of I want to say like 40 keywords and then we had AI go and extract those keywords and see if they were present on the website and then we did that for everyone and then we figured out what was actually useful or we did it on everyone who had an opportunity so it was a much larger sample size and the

[7:35] results were very similar. I mean, I think the specificity of this version is much higher because we might have had something like I forgot the exact keyword we used, but instead of FDIC, which is very specific and tactical, it was like open an account was the the equivalent that we had. And both of those are consumer finance language, but one is extremely specific, easily testable, and the other is like, oh, that could mean a bunch of different things, but generally it's correlated with companies that that are financial services. So it's yeah the the budget cost is significantly down especially on data. Now what your value is is figuring out what the outcomes that you're optimizing for are in that case we were going optimizing for closed one which what should they prioritize uh in terms of accounts. This is a longer sales cycle type uh product. So top level prioritization was the first priority and now it's okay now which which leads within a company should we actually be targeting? And we can actually just take the exact same approach that we did here and say, "Here are all the leads that were on opportunities that closed, go

[8:37] look at all of their attributes that we have, go find more information about them, maybe ones that come from engineering backgrounds are better fits because they understand the problem." So, we're going to run this exact same process on leads. We had one customer run it on who's most likely to pick up the phone if you cold call them. The outcome is really what you're optimizing for, and that's where the the human expertise comes in of what's the business problem we're trying to solve. and then let Cloud Code do all the work to answer that that question for you and give you a like lightweight model that you don't really need to understand the the underlying math for because you can back test it and say if we had used this for the last 6 months go look at all our data and say what the the value prop would and in this scenario you don't need a BDR team anymore do you I mean one operator really is going to be sufficient for running the full workflow Yes and no. I think it what we've actually seen is you get larger BDR teams because the efficiency of a single BDR is now much higher. They're spending much less time getting this raw

[9:39] information. So they're now on calls more often. They're actually like launching campaigns and coming up with new strategies to break into these accounts. So let's say they were talking to uh 20 accounts per day or contacting 20 accounts per day previously. now they're doing 40 or 80 depending on the size of the business SMB versus enterprise. So one BDR the ROI is actually much higher now. So we actually see all of the customers that we work with higher more because the the like yeah revenue per rep is not much higher. Oh so you don't have um a next step in the workflow then that does um personal copy personalization and automated outbound.

[10:21] No, we do. Um it's all through clock code. So I can say in that case why do you need a larger BDR team though? It's for like the meetings and calls. Um a lot of our customers are are doing cold calls and in industries where like fully automated closes aren't aren't necessarily like good fits for the industry yet. Got it. You're saying you still need a human to come in after a reply comes in to kind of book the the meeting maybe. So we do have customers that got rid of the BD BDR function and just AES do everything now and they're doing more volume. We also have customers that are in like the restaurant space where a fully digital close is is not likely in the near term. People want to talk to a person before and the BDRs are able to move that that along. Does that make sense? Is that answering your question correctly? So you're saying with this this technology at this point of time, you still have industries that will require the human operator to step in outside of the workflow to push the needle to get a demo. Is that like for

[11:26] more enterprise or mid-market offerings that Yeah, I would separate into like fully enterprise and and uh SMB plus mid-market. uh if it's like over 50k and over six months sales cycles uh they're doing a lot of different things. Uh BDRs are still useful uh in terms of like getting information from different places and kind of orchestrating the campaigns and making sure that that uh things are getting done. Uh phone calls are still the most valuable uh tool in in the arsenal uh from our perspective. So there there's value there for the mid-market and smaller fuel sizes depending on like the target customer. I I think it differs. We have a PLG company that removed BDRs from their flow because the win rate was just higher with AE's account executives just getting qualified leads from the PLG flow as well as this outbound automated flow. So if your ACV is still low but it's like a AE friendly close that that's a good model. You don't really need a BDR for that. And then we have like high volume, low ACV, big TAM, all

[12:31] restaurants for example. And those folks are actually doing more more touches, more phone calls because the win rate on those is so much higher. So the the value is is so much higher. So yeah, I think you nailed it. It's kind of like dependent on the industry. I might not like talking to a rep to buy software. In fact, I don't and and I'm not the one that is going to be responding to cold calls. Uh but they I like empirically we are seeing cold calls are still extremely effective uh especially in industries that I think I didn't expect them to be. We have a company that we've worked with in like the data analytics data visualization space. So they're selling to like chief data officers or VPs of data and cold calls still work. So I I uh think it it really depends on the type of customer you're you're working with. And do you think do you think the cold calls will at some point be replaced with AI an AI avatar of a BDR?

[13:24] There will be some cases. I think it again goes back to like who's your buyer? A restaurant owner who might not trust that the AI agent is actually like one, do they know if it's an AI agent or not? Do they care? Yeah, they're the videos are getting good. But I think what I would say is like I'm I'm almost going like maybe it's the like optimist in me, but I think relationships are mattering more in specific industries. Like if if I want to buy your product and I already know exactly what I want and I know what I'm doing. We're seeing this a lot in like the developer tooling space where cloud code just makes a recommendation and that's the database we're going to go with. I'm not going to spend a bunch of time. I don't want to talk to a sales rep. when we get to a scale that it matters and we need an enterprise contract, we'll we'll talk. But to start, I I don't want to talk to someone. And so any automated flow will work for me. We don't see that in every industry. Larger contracts are always going to have relationships and relationships are at the end of the day like what you're building here. So there are transactional businesses where you

[14:27] can get away with it and AI agents will take those over for sure. Um I think the other types of businesses or relationship heavy ones will be where we see a lot of uh the value. Cool. And in the background here it's creating a limbless campaign. So you can have that in the skill as well. We try to do a lot of uh human in the loop structure of make sure that you're able to confirm something before the the agent does anything. And I strongly recommend everybody start there just to build trust in the system. But yeah, so we make sure that it doesn't start with a sender attached. You get the lead list in here. It's pre-qualified. So this was just generated from uh the campaign that I added in in the CLI and it comes up with the sequence as well. Perfect.

[15:15] Yeah. So it can go end to end. It's kind of on you to see where you want it to step in, where you're spending time. And the the systems are good enough to where they can do every every step of the flow now. And it's it's on the engineer GTM engineers, the RevOps team to figure out what's the biggest win where we can add the most value. Cool. Anything else that would be helpful? Any kind of results you can uh talk us through? Yeah. Uh we have like the two X outbound campaign using our outbound signals versus the company's existing signals. That was in one recent one was in like the CPG space. what they were using is going to like retail websites and finding brands that were listed, which is relatively constrained. We flipped it around and then the system went and found every company that had a store locator on their website and then ranked and scored them by which stores were in the uh store locator. So, very specific to them, wouldn't make sense for anybody else. And we were getting about 2x better reply rates. So their outbound and LinkedIn campaigns significantly improved. For the larger enterprise

[16:18] companies, the same scoring model for two companies was between 15 and 17% higher close rates. So they were spending more time on on good accounts that ended up closing at a much higher rate. Before working with us, they were spending about 44% of their manh hours or like opportunity time on bad fit accounts uh because they didn't have those negative signals. So they were spending time uh 44% of their time uh for one of the companies uh on things that that were very unlikely to close but because of negative signals that they didn't really notice in real time. Uh and it had to do with like tech stack and staffing. Uh and we surfaced those move those and now they spend more time on good accounts a 17% higher win rate.

[17:02] This is a piece most teams miss. The signal stops being a signal the moment everyone else is pulling from the same three sources. The edge is in the sources nobody has indexed yet and the work is in knowing where to look and how to structure what you find. Jai, thank you for walking through the build links in the description. If you're building in the signal layer, the full breakdown is on GTM Bolt. It maps the episode back to the revenue architecture with the specific patterns worth copying and the ones that depend on your stack. Paid subscribers get the architecture breakdowns in the operator playbooks. Until next time. Thanks, Rick.