{"slug": "how-i-built-an-autonomous-b2b-lead-enrichment-smtp-verification-engine-in-python", "title": "How I Built an Autonomous B2B Lead Enrichment & SMTP Verification Engine in Python for $0.003/Lead", "summary": "A developer built an asynchronous Python lead-enrichment and email-verification engine that combines DNS/MX lookups, raw SMTP handshake checks, concurrent aiohttp scraping of company sites, and a dual-pass LLM pipeline to generate structured outreach icebreakers at roughly $0.003 per lead. The system verifies mailboxes in real time by reading the SMTP response code after RCPT TO rather than sending message bodies, and extracts high-signal page text for LLM processing.", "body_md": "If you have ever scaled outbound campaigns beyond a few hundred contacts a week, you know the exact failure mode: commercial lead databases are stale, bulk email validators miss catch-all domains, and generic \"*Hey {{FirstName}}, loved your profile*\" templates land straight in the spam folder.\n\nMost B2B automation tools force you into an expensive stack ($200+/month for enrichment, $90/month for verification, $150/month for sending tools) that still leaves you with a 4% bounce rate and burned sending domains.\n\nIn this tutorial, we will build a production-grade, asynchronous engine using **Python, `asyncio`, `aiohttp`, DNS/SMTP verification, and an LLM dual-pass pipeline** that:\n\n`/about` or `/blog` pages.\nCommercial enrichment APIs (like Apollo, ZoomInfo, or Clearbit) suffer from two structural problems:\n\nTo solve this, we need **real-time verification** at execution time and **live website extraction** rather than database lookups.\n\nThe pipeline operates across four decoupled asynchronous stages:\n\n```\n[ Raw Lead Input ] \n        │\n        ▼\n[ 1. DNS / MX Lookup & SMTP Handshake Simulation ] ──> (Invalid? Drop/Flag)\n        │ (Valid)\n        ▼\n[ 2. Concurrent Async Web Scraper (aiohttp) ] ──> Fetches /, /about, /blog\n        │\n        ▼\n[ 3. Signal Extraction & HTML Sanitization ] ──> Clean markdown/plain text\n        │\n        ▼\n[ 4. Dual-Pass Structured LLM Engine ] ──> Strict JSON Icebreaker Payload\n        │\n        ▼\n[ Outbound CRM / Webhook Dispatch ]\n```\n\n`aiodns`, establishes a socket connection to the mail server on port 25, sends `HELO` and `MAIL FROM`, followed by `RCPT TO`, and terminates the socket prior to issuing `DATA`.` mistral` or `llama3`) running with `response_format={\"type\": \"json_object\"}`.\nWe avoid triggering spam filters by never sending the actual message body. We only evaluate the SMTP response code after `RCPT TO`:\n\n``` python\nimport asyncio\nimport aiodns\nimport smtplib\nfrom typing import Tuple\n\nasync def get_mx_record(domain: str) -> str:\n    resolver = aiodns.DNSResolver()\n    try:\n        records = await resolver.query(domain, 'MX')\n        records.sort(key=lambda x: x.priority)\n        return str(records[0].host)\n    except Exception:\n        return None\n\nasync def verify_smtp_mailbox(email: str, sender_domain: str = \"verify.yourdomain.com\") -> Tuple[bool, str]:\n    domain = email.split(\"@\")[1]\n    mx_host = await get_mx_record(domain)\n\n    if not mx_host:\n        return False, \"NO_MX_RECORD\"\n\n    # Run SMTP handshake in an async executor to avoid blocking the event loop\n    loop = asyncio.get_event_loop()\n    return await loop.run_in_executor(None, _raw_smtp_handshake, mx_host, email, sender_domain)\n\ndef _raw_smtp_handshake(mx_host: str, target_email: str, sender_domain: str) -> Tuple[bool, str]:\n    try:\n        server = smtplib.SMTP(timeout=10)\n        server.connect(mx_host, 25)\n        server.helo(sender_domain)\n        server.mail(f\"test@{sender_domain}\")\n        code, message = server.rcpt(target_email)\n        server.quit()\n\n        # 250: Requested mail action okay, completed\n        # 550: Requested action not taken: mailbox unavailable\n        if code == 250:\n            return True, \"DELIVERABLE\"\n        elif code == 550:\n            return False, \"MAILBOX_NOT_FOUND\"\n        else:\n            return False, f\"SMTP_CODE_{code}\"\n    except Exception as e:\n        return False, f\"ERROR: {str(e)}\"\n```\n\n`aiohttp` and `BeautifulSoup`\nNext, extract text from the prospect's company URL. We only retain high-signal text elements, discarding scripts, styles, navigations, and footers:\n\n``` python\nimport aiohttp\nfrom bs4 import BeautifulSoup\n\nasync def extract_site_signals(session: aiohttp.ClientSession, base_url: str) -> str:\n    headers = {\"User-Agent\": \"Mozilla/5.0 (compatible; LeadResearchBot/1.0; +https://example.com/bot)\"}\n    try:\n        async with session.get(base_url, headers=headers, timeout=aiohttp.ClientTimeout(total=8)) as response:\n            if response.status != 200:\n                return \"\"\n            html = await response.text()\n\n        soup = BeautifulSoup(html, 'html.parser')\n        for tag in soup([\"script\", \"style\", \"nav\", \"footer\", \"svg\", \"noscript\"]):\n            tag.decompose()\n\n        text = \" \".join(soup.stripped_strings)\n        # Truncate to first 2500 characters to stay within low context token budgets\n        return text[:2500]\n    except Exception:\n        return \"\"\n```\n\nNow we take the extracted plain text and run it through an LLM to generate a bespoke cold outreach opening line. Notice the anti-fluff instructions:\n\n``` python\nfrom openai import AsyncOpenAI\nimport json\n\nclient = AsyncOpenAI(api_key=\"YOUR_OPENAI_API_KEY\")\n\nSYSTEM_PROMPT = \"\"\"\nYou are an expert B2B SDR writing outbound communications.\nExtract a specific technical achievement, recent initiative, or core value proposition from the company text.\nThen write a personalized first-line hook.\n\nRULES:\n- Never say: 'Hope this finds you well', 'I came across your website', or 'I was impressed by'.\n- Point directly to a specific fact.\n- Output MUST conform to the exact JSON schema provided.\n\"\"\"\n\nasync def generate_cold_hook(company_name: str, site_context: str) -> dict:\n    prompt = f\"\"\"\n    Target Company: {company_name}\n    Scraped Context: {site_context}\n\n    Generate the output adhering strictly to this JSON format:\n    {\n      \"detected_core_offering\": \"<string>\",\n      \"specific_detail_noticed\": \"<string>\",\n      \"first_line_icebreaker\": \"<string>\"\n    }\n    \"\"\"\n\n    response = await client.chat.completions.create(\n        model=\"gpt-4o-mini\", # Or your local Ollama proxy\n        response_format={\"type\": \"json_object\"},\n        messages=[\n            {\"role\": \"system\", \"content\": SYSTEM_PROMPT},\n            {\"role\": \"user\", \"content\": prompt}\n        ],\n        temperature=0.4\n    )\n\n    return json.loads(response.choices[0].message.content)\n```\n\nWhen running this at scale, mail servers will drop your IP if you open 50 connections at once to the same host (e.g., Google or Outlook).\n\nImplement an `asyncio.Semaphore` along with domain grouping:\n\n``` python\nasync def process_lead(sem: asyncio.Semaphore, session: aiohttp.ClientSession, lead: dict):\n    async with sem:\n        # 1. SMTP Check\n        is_valid, reason = await verify_smtp_mailbox(lead[\"email\"])\n        if not is_valid:\n            return {**lead, \"status\": \"FAILED_VERIFICATION\", \"reason\": reason}\n\n        # 2. Extract Signals\n        context = await extract_site_signals(session, lead[\"website\"])\n        if not context:\n            return {**lead, \"status\": \"FAILED_SCRAPING\"}\n\n        # 3. LLM Synthesis\n        enrichment = await generate_cold_hook(lead[\"company_name\"], context)\n\n        return {\n            **lead,\n            \"status\": \"READY\",\n            **enrichment\n        }\n\nasync def main():\n    concurrency_limit = asyncio.Semaphore(15) # Safe boundary\n    async with aiohttp.ClientSession() as session:\n        tasks = [process_lead(concurrency_limit, session, lead) for lead in leads_list]\n        results = await asyncio.gather(*tasks)\n        # Write results to database or push to CRM webhook\n```\n\n`gpt-4o-mini` tokens; $0.000 if using a local Ollama instance like `mistral:7b-instruct`).\nYou now have the blueprint to run high-volume, low-cost outbound pipelines without relying on brittle third-party providers. By combining direct async DNS/SMTP checks with targeted extraction, your domain reputation stays pristine while your cold email response rates climb.\n\nYou can implement this architecture by copying the snippets above into your custom backend, or grab the fully configured, production-tested engine complete with test fixtures, automated CRM push webhooks, and n8n integration schemas:\n\n`EARLYBIRD` at checkout for 20% off).\nDrop a comment below if you have questions about handling MX catch-all responses or fine-tuning local Ollama models for high-concurrency parsing!", "url": "https://wpnews.pro/news/how-i-built-an-autonomous-b2b-lead-enrichment-smtp-verification-engine-in-python", "canonical_source": "https://dev.to/reigen/how-i-built-an-autonomous-b2b-lead-enrichment-smtp-verification-engine-in-python-for-0003lead-1cga", "published_at": "2026-09-30 18:37:14+00:00", "updated_at": "2026-09-30 18:46:42.632885+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "developer-tools"], "entities": ["Python", "aiohttp", "aiodns", "BeautifulSoup", "Apollo", "ZoomInfo", "Clearbit", "Mistral"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-built-an-autonomous-b2b-lead-enrichment-smtp-verification-engine-in-python", "markdown": "https://wpnews.pro/news/how-i-built-an-autonomous-b2b-lead-enrichment-smtp-verification-engine-in-python.md", "text": "https://wpnews.pro/news/how-i-built-an-autonomous-b2b-lead-enrichment-smtp-verification-engine-in-python.txt", "jsonld": "https://wpnews.pro/news/how-i-built-an-autonomous-b2b-lead-enrichment-smtp-verification-engine-in-python.jsonld"}}