{"slug": "engineering-agentic-systems-for-financial-workflows-harnessing-llm-non-while", "title": "Engineering Agentic Systems for Financial Workflows: Harnessing LLM Non-Determinism While Guaranteeing Deterministic Execution", "summary": "An engineering writeup describes an 80/20 hybrid architecture for financial workflows in which LLMs act only as a cognitive coordinator that interprets unstructured inputs and emits typed specifications, while deterministic, Pydantic-typed skill functions such as variance_analysis() and invoice_match() handle all arithmetic, ledger balancing, and payment authorization. The piece argues LLMs must never serve as the primary execution engine or system of record for double-entry bookkeeping, tax and FX conversions, or database writes.", "body_md": "The integration of Large Language Models (LLMs) into financial technology introduces a fundamental engineering challenge: **financial systems require absolute mathematical determinism, whereas LLMs are intrinsically probabilistic and non-deterministic engines**. \n\nWhen financial engineering teams attempt to use LLMs as direct calculation engines, failure is inevitable. LLMs are notoriously ill-suited for arithmetic, ledger balancing, interest rate calculations, tax computations, FX conversions, and exact reconciliation rules. However, attempting to eliminate non-determinism entirely by restricting enterprise automation to hardcoded rule engines leaves institutions incapable of processing the unstructured, ambiguous, and multi-source realities of modern commerce.\n\nThe core architectural breakthrough lies in shifting the paradigm: **the objective is not to make the LLM compute financial numbers, but to utilize LLM non-determinism as a cognitive coordinator that interprets unstructured ambiguity, forms hypotheses, and emits typed specifications for deterministic engines to execute**.\n\n## \n  \n  \n  1. The Core Selection Test: Where Non-Determinism Adds Value\n\nTo build a reliable agentic financial system, architects must apply a rigorous evaluation framework before assigning any task to an LLM.\n\n### \n  \n  \n  When Non-Determinism is Essential\n\nA financial workflow is an optimal candidate for an LLM when it satisfies the following six criteria:\n\n1. \n**Input Variability:** The task originates from unstructured or semi-structured sources—such as natural language emails, PDF invoices, free-text remittance notes, customer support tickets, or inconsistent source schemas.\n2. \n**Path Uncertainty:** The required sequence of API calls or database queries cannot be pre-calculated before inspecting the case details.\n3. \n**Semantic Ambiguity:** Contextual evaluation is required to interpret ambiguous business terminology (e.g., distinguishing whether a billing variation is a \"temporary migration overlap,\" a \"contractual discount,\" or an \"unauthorized discrepancy\").\n4. \n**Cross-Source Synthesis:** Resolving the case requires correlating structured general ledger (GL) entries with unstructured documentation, such as procurement contracts, customer threads, or external vendor policies.\n5. \n**Multiple Plausible Hypotheses:** The system must generate and systematically test competing candidate explanations before evidence rules them out.\n6. \n**Human-Tailored Explanation:** The final output requires translating verified numerical results into contextual business narratives tailored to specific stakeholders (e.g., CFOs, collections managers, or end customers).\n\n### \n  \n  \n  Tasks That Must Remain 100% Deterministic\n\nLLMs must **never** operate as the primary execution engine or system of record for:\n\n- Double-entry bookkeeping and trial balance validation.\n- Arithmetic, interest, tax, and FX conversions.\n- Revenue and EBITDA formula calculations.\n- Database writes, payment authorizations, and spending limit enforcement.\n- State machine transitions, access control, and transaction permissions.\n\n## \n  \n  \n  2. The 80/20 Layered Technical Architecture\n\nTo achieve sub-second execution safety and zero-hallucination guarantees, financial architectures employ an **80/20 hybrid layered model**.\n\n### \n  \n  \n  Layer 1: Preconfigured Financial Skills (80% of System Logic)\n\nThis layer consists of deterministic code functions with strictly typed Pydantic schemas, idempotent execution guarantees, unit test coverage, and audit logging. Key skills include:\n\n- `variance_analysis()`\n- `price_volume_mix()`\n- `invoice_match()`\n- `payment_candidate_match()`\n- `aging_analysis()`\n- `claim_validate()`\n\n### \n  \n  \n  Layer 2: Dynamic Reasoning Layer (20% of System Logic)\n\nThe LLM sits in this layer as the orchestration engine. It evaluates incoming unstructured events, determines which Layer 1 skills to invoke, identifies missing context, evaluates returned evidence, and constructs candidate resolution proposals.\n\n### \n  \n  \n  Layer 3: Custom Calculation Sandbox\n\nFor complex, ad-hoc scenario modeling that falls outside preconfigured skills, the LLM must **not** execute arbitrary code or shell scripts. Instead, it outputs a constrained, typed `CalculationSpec`:\n\nThis specification is validated against Pydantic schemas and policy constraints before execution in an isolated Python runtime, returning a fully deterministic, auditable execution trace.\n\n## \n  \n  \n  3. High-Value Implementation Case Studies\n\n### \n  \n  \n  Case Study A: Payment & Reconciliation Exception Resolution\n\nA common operational bottleneck occurs in cash application when incoming payments cannot be matched directly to open receivables.\n\n#### \n  \n  \n  The Problem Context\n\n- \n**Bank Stream Event:**`ACME TECH SERVICES 142,750 USD`\n- \n**Internal Records:** Invoice A (`$142,750` ), Invoice B (`$71,375` ), Credit Note (`$71,375` ).\n- \n**Customer Communication:***\"We settled both open balances after applying the agreed adjustment.\"*\n- \n**Rule Engine Failure:** Traditional string-matching or exact-amount lookups fail because the transaction descriptor does not reference a single invoice number.\n\n#### \n  \n  \n  Orchestrated Division of Labor\n\n- \n**LLM Role:** Interprets unstructured remittance text, parses email threads, correlates cross-source records, formulates allocation hypotheses, and drafts resolution proposals.\n- \n**Deterministic Code Role:** Queries ledger database, calculates residuals, executes exact-match mathematical verifications, enforces double-allocation prevention, and gates ledger postings behind policy thresholds or human approvals.\n\n### \n  \n  \n  Case Study B: Serverless Real-Time KYC Architecture on AWS\n\nModernizing Know Your Customer (KYC) workflows requires transitioning from legacy, batch-oriented monolithic architectures (which average 3–5 days per onboarding case) to event-driven, real-time agentic pipelines.\n\n#### \n  \n  \n  Architectural Breakdown\n\n1. \n**Event Streaming Backbone:****Amazon Managed Streaming for Apache Kafka (Amazon MSK)** ingests inbound customer requests, document uploads, and third-party verification events asynchronously.\n2. \n**Orchestration Runtime:****Amazon Bedrock AgentCore** hosts the runtime environment, providing native session state persistence, context management, and multi-agent coordination.\n3. \n**Supervisor & Domain-Specific Sub-Agents:** The**KYC Orchestration Supervisor Agent** dynamically constructs execution plans across five specialized sub-agents:  - \n**Identity Verification Sub-Agent:** Cross-references customer identities against global watchlists, PEP, and sanctions databases.\n  - \n**Document Analysis Sub-Agent:** Performs OCR on identity documents, assesses image quality, translates multi-language documents, and detects document forgery.\n  - \n**Fraud Detection Sub-Agent:** Conducts behavioral analysis, evaluates IP duplicate applications, and performs semantic similarity searches over historical fraud vectors.\n  - \n**Compliance & Risk Sub-Agent:** Interprets jurisdiction-specific regulations (e.g., BSA, USA PATRIOT Act, EU AMLD, MAS guidelines) and generates verifiable audit attestations.\n  - \n**Customer Experience Sub-Agent:** Identifies onboarding friction points and optimizes applicant communication.\n4. \n**Retrieval-Augmented Generation (RAG):****Amazon OpenSearch Serverless** provides vector search over institutional policies stored in**Amazon S3** , while**Amazon DynamoDB** provides sub-millisecond status lookups.\n5. \n**Dynamic Confidence Routing:** The Supervisor Agent routes outcomes based on sub-agent confidence scores:  - \n**High Confidence (>95%):** Automated real-time approval.\n  - \n**Medium Confidence (75%–95%):** Triggers step-up verification workflows.\n  - \n**Low Confidence (<75%):** Escalates to human compliance reviewers with an auto-generated decision packet.\n\n#### \n  \n  \n  Performance & Operational ROI\n\n- \n**Validation Latency:** Reduced from**3–5 days to under 5 minutes** .\n- \n**Compliance Efficiency:** Automated routing enables compliance specialists to handle**up to 4x their previous caseload** by focusing exclusively on complex, escalated exceptions.\n\n## \n  \n  \n  4. Strategic Portfolio Analysis of Financial Agentic Workflows\n\nWhen prioritizing agentic implementations across financial operations, workflows are evaluated by business impact, money-handling risks, and architectural feasibility:\n\n| Workflow Rank & Title | Core LLM Functionality | Deterministic System Functionality | Business Impact & Primary Metrics | \n| **Rank 1: Payment Reconciliation Exception Agent** | Unstructured remittance parsing, cross-source entity resolution, hypothesis generation. | Residual arithmetic, candidate matching, double-allocation constraints, ledger write gates. | **High ROI:** Dramatically increases automated cash application rate; cuts analyst resolution time per exception. | \n| **Rank 2: Accounts Receivable (AR) Case Handling Agent** | Communication analysis, dispute classification, automated contextual response drafting. | Aging calculation, outstanding balance computation, task scheduling, policy enforcement. | **Direct Cash Impact:** Decreases Days Sales Outstanding (DSO) and reduces manual collector overhead. | \n| **Rank 3: Variance-to-Decision FP&A Agent** | Translating business queries into investigation paths, hypothesis generation over EBITDA drivers. | Bridge calculations, price-volume-mix formulas, period/currency conversions. | **Decision Velocity:** Compresses monthly variance analysis cycles from days to minutes. | \n| **Rank 4: Cash-Forecast Exception Investigator** | Identifying operational drivers behind forecast deviations, qualitative impact synthesis. | Cash runway modeling, scenario impact quantification, threshold monitoring. | **Risk Reduction:** Accelerates early warning detection for liquidity risks. | \n| **Auxiliary: Management Intent to Analysis Plan** | Parsing vague prompts (e.g., *\"Can we afford 3 hires if sales slow?\"* ) into typed`AnalysisRequest` objects. | Executing financial scenario projections, cash constraint validations. | **UX Transformation:** Bridges non-technical business intent with complex modeling engines. | \n| **Auxiliary: Case Escalation Packet Generator** | Synthesizing investigated evidence, failed hypotheses, and citations into compact decision summaries. | Evidence package bundling, permission checks, audit logging. | **Handoff Efficiency:** Eliminates context reconstruction time during human escalations. | \n\n## \n  \n  \n  5. Architectural Principles for Engineering Agentic Finance\n\n1. \n**Enforce Rigid Separation Between Reasoning and Calculation:** Never permit an LLM to directly calculate monetary figures or perform ledger state transitions. The LLM generates structured investigation plans; code executes the operations and validates the math.\n2. \n**Implement Claim Schema Verification:** Before presenting an LLM-synthesized narrative to a user or reviewer, run a verifier component that validates every numerical claim against underlying source database records.\n3. \n**Design for Fallback and Human Escalation:** Establish explicit confidence thresholds. High-risk or low-confidence actions must emit standardized decision packets for human approval before execution.\n4. \n**Scope Tasks Around Complete Units of Work:** Build agents around bounded operational problems that feature a structured trigger, an unstructured middle, deterministic tool availability, a safe action boundary, and clear evaluation metrics.", "url": "https://wpnews.pro/news/engineering-agentic-systems-for-financial-workflows-harnessing-llm-non-while", "canonical_source": "https://dev.to/_aparna_pradhan_/engineering-agentic-systems-for-financial-workflows-harnessing-llm-non-determinism-while-5acm", "published_at": "2026-09-30 04:36:47+00:00", "updated_at": "2026-09-30 04:46:41.998688+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "mlops"], "entities": ["Pydantic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/engineering-agentic-systems-for-financial-workflows-harnessing-llm-non-while", "markdown": "https://wpnews.pro/news/engineering-agentic-systems-for-financial-workflows-harnessing-llm-non-while.md", "text": "https://wpnews.pro/news/engineering-agentic-systems-for-financial-workflows-harnessing-llm-non-while.txt", "jsonld": "https://wpnews.pro/news/engineering-agentic-systems-for-financial-workflows-harnessing-llm-non-while.jsonld"}}