I built a Creator Engine that extracts promises from sales conversations (80% accuracy) A developer built the Creator Engine, an LLM-based system that compares promises made in sales conversations against actual contract terms, reporting roughly 80% accuracy on a 20-case Promise Alignment Benchmark v0. The architecture follows a strict rule that the LLM only extracts while deterministic engine-level invariant checks decide outcomes, yielding zero invariant violations across all benchmark cases after three prompt iterations. The project is published as a benchmark repo on GitHub and is intended to power five planned products sharing a common Comparison Intelligence core. Last week I built 32 MCP servers. Today I want to share the more ambitious project: the Creator Engine — a system that compares what sales promises to what the contract actually says. Sales teams promise things. Contracts specify things. They don't always match. Each mismatch is potential revenue loss, legal exposure, or customer churn. Traditional CLM tools are expensive $50K+/yr , heavy, and focus on drafts — not on reality alignment. Promise Alignment Benchmark v0 — an executable specification for testing the engine. creator-engine/benchmarks/promise-alignment/v0/ contains: Cover: HARD FAILURE — 19/20 = FAIL if any invariant breaks. Three-step LLM pipeline: Then engine-level invariant checks run on top. For example: function runInvariants promises, terms, alignments { const violations = ; for const p of promises { if p.status === "NOT COMMITMENT" { violations.push { id: "INV-004", reason: "NOT COMMITMENT promise" } ; } } for const a of alignments { if a.relationship === "CONFLICT" && a.reasoning { violations.push { id: "INV-001", reason: "CONFLICT without reasoning" } ; } if a.impact && a.impact.amount = null && a.impact.calculation id { violations.push { id: "INV-002", reason: "amount without calculation id" } ; } } return violations; } Architectural rule: LLM extracts, engine decides. Never LLM to final truth. After 3 prompt iterations: Zero invariant violations across all 20 cases. Failures that remain are mostly edge cases where ground truth is debatable: This is expected. 80% is a baseline, not production. Next: multi-step pipeline + verification layer Phase 42 . GitHub: Alex-Dev-Web-C-DE/creator-engine Not on npm yet — it's a benchmark repo, not a package. The engine will power 5 products: Sales Promise Tracker, ScopeGuard, InvoiceGuard, DeliveryGuard, RevenueLeak. All share the same core: Comparison Intelligence. While the engine matures, you can try the live MCP servers: 35 servers on MCPize — validators, parsers, extractors. Most relevant for the engine use case: All x402: agents pay USDC on Base, no signup, no API keys. What's your take? Is 80% good enough to ship a human-in-the-loop MVP, or should I push for 90%+ before any customer touches it?