cd /news/ai-agents/engineering-agentic-systems-for-fina… · home › topics › ai-agents › article
[ARTICLE · art-142267] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Engineering Agentic Systems for Financial Workflows: Harnessing LLM Non-Determinism While Guaranteeing Deterministic Execution

An engineering writeup describes an 80/20 hybrid architecture for financial workflows in which LLMs act only as a cognitive coordinator that interprets unstructured inputs and emits typed specifications, while deterministic, Pydantic-typed skill functions such as variance_analysis() and invoice_match() handle all arithmetic, ledger balancing, and payment authorization. The piece argues LLMs must never serve as the primary execution engine or system of record for double-entry bookkeeping, tax and FX conversions, or database writes.

by read7 min views1 publishedSep 30, 2026

The integration of Large Language Models (LLMs) into financial technology introduces a fundamental engineering challenge: financial systems require absolute mathematical determinism, whereas LLMs are intrinsically probabilistic and non-deterministic engines.

When financial engineering teams attempt to use LLMs as direct calculation engines, failure is inevitable. LLMs are notoriously ill-suited for arithmetic, ledger balancing, interest rate calculations, tax computations, FX conversions, and exact reconciliation rules. However, attempting to eliminate non-determinism entirely by restricting enterprise automation to hardcoded rule engines leaves institutions incapable of processing the unstructured, ambiguous, and multi-source realities of modern commerce.

The core architectural breakthrough lies in shifting the paradigm: the objective is not to make the LLM compute financial numbers, but to utilize LLM non-determinism as a cognitive coordinator that interprets unstructured ambiguity, forms hypotheses, and emits typed specifications for deterministic engines to execute.

#

  1. The Core Selection Test: Where Non-Determinism Adds Value

To build a reliable agentic financial system, architects must apply a rigorous evaluation framework before assigning any task to an LLM.

When Non-Determinism is Essential

A financial workflow is an optimal candidate for an LLM when it satisfies the following six criteria:

Input Variability: The task originates from unstructured or semi-structured sources—such as natural language emails, PDF invoices, free-text remittance notes, customer support tickets, or inconsistent source schemas. 2. Path Uncertainty: The required sequence of API calls or database queries cannot be pre-calculated before inspecting the case details. 3. Semantic Ambiguity: Contextual evaluation is required to interpret ambiguous business terminology (e.g., distinguishing whether a billing variation is a "temporary migration overlap," a "contractual discount," or an "unauthorized discrepancy"). 4. Cross-Source Synthesis: Resolving the case requires correlating structured general ledger (GL) entries with unstructured documentation, such as procurement contracts, customer threads, or external vendor policies. 5. Multiple Plausible Hypotheses: The system must generate and systematically test competing candidate explanations before evidence rules them out. 6. Human-Tailored Explanation: The final output requires translating verified numerical results into contextual business narratives tailored to specific stakeholders (e.g., CFOs, collections managers, or end customers).

Tasks That Must Remain 100% Deterministic

LLMs must never operate as the primary execution engine or system of record for:

  • Double-entry bookkeeping and trial balance validation.
  • Arithmetic, interest, tax, and FX conversions.
  • Revenue and EBITDA formula calculations.
  • Database writes, payment authorizations, and spending limit enforcement.
  • State machine transitions, access control, and transaction permissions.

#

  1. The 80/20 Layered Technical Architecture

To achieve sub-second execution safety and zero-hallucination guarantees, financial architectures employ an 80/20 hybrid layered model.

Layer 1: Preconfigured Financial Skills (80% of System Logic)

This layer consists of deterministic code functions with strictly typed Pydantic schemas, idempotent execution guarantees, unit test coverage, and audit logging. Key skills include:

- `variance_analysis()`
- `price_volume_mix()`
- `invoice_match()`
- `payment_candidate_match()`
- `aging_analysis()`
- `claim_validate()`

Layer 2: Dynamic Reasoning Layer (20% of System Logic) The LLM sits in this layer as the orchestration engine. It evaluates incoming unstructured events, determines which Layer 1 skills to invoke, identifies missing context, evaluates returned evidence, and constructs candidate resolution proposals.

Layer 3: Custom Calculation Sandbox

For complex, ad-hoc scenario modeling that falls outside preconfigured skills, the LLM must not execute arbitrary code or shell scripts. Instead, it outputs a constrained, typed CalculationSpec: This specification is validated against Pydantic schemas and policy constraints before execution in an isolated Python runtime, returning a fully deterministic, auditable execution trace.

#

  1. High-Value Implementation Case Studies

Case Study A: Payment & Reconciliation Exception Resolution A common operational bottleneck occurs in cash application when incoming payments cannot be matched directly to open receivables.

The Problem Context

Bank Stream Event:ACME TECH SERVICES 142,750 USD #

Internal Records: Invoice A ($142,750 ), Invoice B ($71,375 ), Credit Note ($71,375 ). #

Customer Communication:"We settled both open balances after applying the agreed adjustment." #

Rule Engine Failure: Traditional string-matching or exact-amount lookups fail because the transaction descriptor does not reference a single invoice number.

Orchestrated Division of Labor

LLM Role: Interprets unstructured remittance text, parses email threads, correlates cross-source records, formulates allocation hypotheses, and drafts resolution proposals. #

Deterministic Code Role: Queries ledger database, calculates residuals, executes exact-match mathematical verifications, enforces double-allocation prevention, and gates ledger postings behind policy thresholds or human approvals.

Case Study B: Serverless Real-Time KYC Architecture on AWS Modernizing Know Your Customer (KYC) workflows requires transitioning from legacy, batch-oriented monolithic architectures (which average 3–5 days per onboarding case) to event-driven, real-time agentic pipelines.

Architectural Breakdown

**Event Streaming Backbone:Amazon Managed Streaming for Apache Kafka (Amazon MSK) ingests inbound customer requests, document uploads, and third-party verification events asynchronously. 2. Orchestration Runtime:**Amazon Bedrock AgentCore hosts the runtime environment, providing native session state persistence, context management, and multi-agent coordination. 3. Supervisor & Domain-Specific Sub-Agents: TheKYC Orchestration Supervisor Agent dynamically constructs execution plans across five specialized sub-agents: - Identity Verification Sub-Agent: Cross-references customer identities against global watchlists, PEP, and sanctions databases. #

Document Analysis Sub-Agent: Performs OCR on identity documents, assesses image quality, translates multi-language documents, and detects document forgery. #

Fraud Detection Sub-Agent: Conducts behavioral analysis, evaluates IP duplicate applications, and performs semantic similarity searches over historical fraud vectors. #

Compliance & Risk Sub-Agent: Interprets jurisdiction-specific regulations (e.g., BSA, USA PATRIOT Act, EU AMLD, MAS guidelines) and generates verifiable audit attestations. #

Customer Experience Sub-Agent: Identifies onboarding friction points and optimizes applicant communication. 4. Retrieval-Augmented Generation (RAG):**Amazon OpenSearch Serverless provides vector search over institutional policies stored inAmazon S3** , whileAmazon DynamoDB provides sub-millisecond status lookups. 5. Dynamic Confidence Routing: The Supervisor Agent routes outcomes based on sub-agent confidence scores: - High Confidence (>95%): Automated real-time approval. #

Medium Confidence (75%–95%): Triggers step-up verification workflows. #

Low Confidence (<75%): Escalates to human compliance reviewers with an auto-generated decision packet.

Performance & Operational ROI

Validation Latency: Reduced from3–5 days to under 5 minutes . #

Compliance Efficiency: Automated routing enables compliance specialists to handleup to 4x their previous caseload by focusing exclusively on complex, escalated exceptions.

#

  1. Strategic Portfolio Analysis of Financial Agentic Workflows

When prioritizing agentic implementations across financial operations, workflows are evaluated by business impact, money-handling risks, and architectural feasibility:

| Workflow Rank & Title | Core LLM Functionality | Deterministic System Functionality | Business Impact & Primary Metrics | | Rank 1: Payment Reconciliation Exception Agent | Unstructured remittance parsing, cross-source entity resolution, hypothesis generation. | Residual arithmetic, candidate matching, double-allocation constraints, ledger write gates. | High ROI: Dramatically increases automated cash application rate; cuts analyst resolution time per exception. | | Rank 2: Accounts Receivable (AR) Case Handling Agent | Communication analysis, dispute classification, automated contextual response drafting. | Aging calculation, outstanding balance computation, task scheduling, policy enforcement. | Direct Cash Impact: Decreases Days Sales Outstanding (DSO) and reduces manual collector overhead. | | Rank 3: Variance-to-Decision FP&A Agent | Translating business queries into investigation paths, hypothesis generation over EBITDA drivers. | Bridge calculations, price-volume-mix formulas, period/currency conversions. | Decision Velocity: Compresses monthly variance analysis cycles from days to minutes. | | Rank 4: Cash-Forecast Exception Investigator | Identifying operational drivers behind forecast deviations, qualitative impact synthesis. | Cash runway modeling, scenario impact quantification, threshold monitoring. | Risk Reduction: Accelerates early warning detection for liquidity risks. | | Auxiliary: Management Intent to Analysis Plan | Parsing vague prompts (e.g., "Can we afford 3 hires if sales slow?" ) into typedAnalysisRequest objects. | Executing financial scenario projections, cash constraint validations. | UX Transformation: Bridges non-technical business intent with complex modeling engines. | | Auxiliary: Case Escalation Packet Generator | Synthesizing investigated evidence, failed hypotheses, and citations into compact decision summaries. | Evidence package bundling, permission checks, audit logging. | Handoff Efficiency: Eliminates context reconstruction time during human escalations. |

#

  1. Architectural Principles for Engineering Agentic Finance

Enforce Rigid Separation Between Reasoning and Calculation: Never permit an LLM to directly calculate monetary figures or perform ledger state transitions. The LLM generates structured investigation plans; code executes the operations and validates the math. 2. Implement Claim Schema Verification: Before presenting an LLM-synthesized narrative to a user or reviewer, run a verifier component that validates every numerical claim against underlying source database records. 3. Design for Fallback and Human Escalation: Establish explicit confidence thresholds. High-risk or low-confidence actions must emit standardized decision packets for human approval before execution. 4. Scope Tasks Around Complete Units of Work: Build agents around bounded operational problems that feature a structured trigger, an unstructured middle, deterministic tool availability, a safe action boundary, and clear evaluation metrics.

── more in #ai-agents 4 stories · sorted by recency
── more on @pydantic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/engineering-agentic-…] indexed:0 read:7min 2026-09-30 · —