How I Built an Autonomous B2B Lead Enrichment & SMTP Verification Engine in Python for $0.003/Lead A developer built an asynchronous Python lead-enrichment and email-verification engine that combines DNS/MX lookups, raw SMTP handshake checks, concurrent aiohttp scraping of company sites, and a dual-pass LLM pipeline to generate structured outreach icebreakers at roughly $0.003 per lead. The system verifies mailboxes in real time by reading the SMTP response code after RCPT TO rather than sending message bodies, and extracts high-signal page text for LLM processing. If you have ever scaled outbound campaigns beyond a few hundred contacts a week, you know the exact failure mode: commercial lead databases are stale, bulk email validators miss catch-all domains, and generic " Hey {{FirstName}}, loved your profile " templates land straight in the spam folder. Most B2B automation tools force you into an expensive stack $200+/month for enrichment, $90/month for verification, $150/month for sending tools that still leaves you with a 4% bounce rate and burned sending domains. In this tutorial, we will build a production-grade, asynchronous engine using Python, asyncio , aiohttp , DNS/SMTP verification, and an LLM dual-pass pipeline that: /about or /blog pages. Commercial enrichment APIs like Apollo, ZoomInfo, or Clearbit suffer from two structural problems: To solve this, we need real-time verification at execution time and live website extraction rather than database lookups. The pipeline operates across four decoupled asynchronous stages: Raw Lead Input │ ▼ 1. DNS / MX Lookup & SMTP Handshake Simulation ── Invalid? Drop/Flag │ Valid ▼ 2. Concurrent Async Web Scraper aiohttp ── Fetches /, /about, /blog │ ▼ 3. Signal Extraction & HTML Sanitization ── Clean markdown/plain text │ ▼ 4. Dual-Pass Structured LLM Engine ── Strict JSON Icebreaker Payload │ ▼ Outbound CRM / Webhook Dispatch aiodns , establishes a socket connection to the mail server on port 25, sends HELO and MAIL FROM , followed by RCPT TO , and terminates the socket prior to issuing DATA . mistral or llama3 running with response format={"type": "json object"} . We avoid triggering spam filters by never sending the actual message body. We only evaluate the SMTP response code after RCPT TO : python import asyncio import aiodns import smtplib from typing import Tuple async def get mx record domain: str - str: resolver = aiodns.DNSResolver try: records = await resolver.query domain, 'MX' records.sort key=lambda x: x.priority return str records 0 .host except Exception: return None async def verify smtp mailbox email: str, sender domain: str = "verify.yourdomain.com" - Tuple bool, str : domain = email.split "@" 1 mx host = await get mx record domain if not mx host: return False, "NO MX RECORD" Run SMTP handshake in an async executor to avoid blocking the event loop loop = asyncio.get event loop return await loop.run in executor None, raw smtp handshake, mx host, email, sender domain def raw smtp handshake mx host: str, target email: str, sender domain: str - Tuple bool, str : try: server = smtplib.SMTP timeout=10 server.connect mx host, 25 server.helo sender domain server.mail f"test@{sender domain}" code, message = server.rcpt target email server.quit 250: Requested mail action okay, completed 550: Requested action not taken: mailbox unavailable if code == 250: return True, "DELIVERABLE" elif code == 550: return False, "MAILBOX NOT FOUND" else: return False, f"SMTP CODE {code}" except Exception as e: return False, f"ERROR: {str e }" aiohttp and BeautifulSoup Next, extract text from the prospect's company URL. We only retain high-signal text elements, discarding scripts, styles, navigations, and footers: python import aiohttp from bs4 import BeautifulSoup async def extract site signals session: aiohttp.ClientSession, base url: str - str: headers = {"User-Agent": "Mozilla/5.0 compatible; LeadResearchBot/1.0; +https://example.com/bot "} try: async with session.get base url, headers=headers, timeout=aiohttp.ClientTimeout total=8 as response: if response.status = 200: return "" html = await response.text soup = BeautifulSoup html, 'html.parser' for tag in soup "script", "style", "nav", "footer", "svg", "noscript" : tag.decompose text = " ".join soup.stripped strings Truncate to first 2500 characters to stay within low context token budgets return text :2500 except Exception: return "" Now we take the extracted plain text and run it through an LLM to generate a bespoke cold outreach opening line. Notice the anti-fluff instructions: python from openai import AsyncOpenAI import json client = AsyncOpenAI api key="YOUR OPENAI API KEY" SYSTEM PROMPT = """ You are an expert B2B SDR writing outbound communications. Extract a specific technical achievement, recent initiative, or core value proposition from the company text. Then write a personalized first-line hook. RULES: - Never say: 'Hope this finds you well', 'I came across your website', or 'I was impressed by'. - Point directly to a specific fact. - Output MUST conform to the exact JSON schema provided. """ async def generate cold hook company name: str, site context: str - dict: prompt = f""" Target Company: {company name} Scraped Context: {site context} Generate the output adhering strictly to this JSON format: { "detected core offering": "