Last week I built 32 MCP servers. Today I want to share the more ambitious project: the Creator Engine — a system that compares what sales promises to what the contract actually says.
Sales teams promise things. Contracts specify things. They don't always match.
Each mismatch is potential revenue loss, legal exposure, or customer churn. Traditional CLM tools are expensive ($50K+/yr), heavy, and focus on drafts — not on reality alignment.
Promise Alignment Benchmark v0 — an executable specification for testing the engine.
creator-engine/benchmarks/promise-alignment/v0/ contains: Cover:
HARD FAILURE — 19/20 = FAIL if any invariant breaks.
Three-step LLM pipeline: Then engine-level invariant checks run on top. For example:
function runInvariants(promises, terms, alignments) {
const violations = [];
for (const p of promises) {
if (p.status === "NOT_COMMITMENT") {
violations.push({ id: "INV-004", reason: "NOT_COMMITMENT promise" });
}
}
for (const a of alignments) {
if (a.relationship === "CONFLICT" && !a.reasoning) {
violations.push({ id: "INV-001", reason: "CONFLICT without reasoning" });
}
if (a.impact && a.impact.amount != null && !a.impact.calculation_id) {
violations.push({ id: "INV-002", reason: "amount without calculation_id" });
}
}
return violations;
}
Architectural rule: LLM extracts, engine decides. Never LLM to final truth.
After 3 prompt iterations:
Zero invariant violations across all 20 cases.
Failures that remain are mostly edge cases where ground truth is debatable:
This is expected. 80% is a baseline, not production. Next: multi-step pipeline + verification layer (Phase 42).
GitHub: Alex-Dev-Web-C-DE/creator-engine Not on npm yet — it's a benchmark repo, not a package.
The engine will power 5 products: Sales Promise Tracker, ScopeGuard, InvoiceGuard, DeliveryGuard, RevenueLeak.
All share the same core: Comparison Intelligence.
While the engine matures, you can try the live MCP servers: 35 servers on MCPize — validators, parsers, extractors. Most relevant for the engine use case:
All x402: agents pay USDC on Base, no signup, no API keys.
What's your take? Is 80% good enough to ship a human-in-the-loop MVP, or should I push for 90%+ before any customer touches it?