# Asynchronous Parallel Validation & Diff Report Generation Tool for Multiple AI Platforms

> Source: <https://dev.to/toai/asynchronous-parallel-validation-diff-report-generation-tool-for-multiple-ai-platforms-2p7n>
> Published: 2026-10-08 00:10:48+00:00

Here is the fully translated and refined English article, tailored for a technical audience on Dev.to. All technical details, structural elements, and the complete code block have been preserved and translated into professional engineering terminology. The conclusion has been distilled to focus purely on the technical takeaways.

To achieve a script that is highly portable, runs immediately in any environment without prior setup, and minimizes dependencies, we adopted the following architectural design:

`urllib.request` for HTTP communications, alongside `json` and `difflib`.` ProcessPoolExecutor`.` as_completed` loop, applying limits via `future.result(timeout=remaining)`.
Below is the complete code that was evaluated during our testing phase—a script that ultimately succumbed to deadlocks and cascading timeouts.

``` python
import sys
import json
import urllib.request
import urllib.error
from concurrent.futures import ProcessPoolExecutor, as_completed
import time
import difflib

def call_endpoint(endpoint_info):
    """
    Sends an HTTP request to a single endpoint and returns the result.
    Placed at the module top-level to allow serialization by the process pool.
    """
    name = endpoint_info.get("name", "Unknown")
    url = endpoint_info.get("url")
    headers = endpoint_info.get("headers", {})
    payload = endpoint_info.get("payload", {})
    timeout = endpoint_info.get("timeout", 10)

    start_time = time.time()
    try:
        data = json.dumps(payload).encode("utf-8")
        req = urllib.request.Request(url, data=data, headers=headers, method="POST")

        with urllib.request.urlopen(req, timeout=timeout) as response:
            elapsed = time.time() - start_time
            body = response.read().decode("utf-8")
            try:
                parsed_body = json.loads(body)
            except json.JSONDecodeError:
                parsed_body = body

            return {
                "name": name,
                "status": "success",
                "status_code": response.status,
                "elapsed": round(elapsed, 3),
                "response": parsed_body
            }
    except urllib.error.HTTPError as e:
        elapsed = time.time() - start_time
        err_body = e.read().decode("utf-8", errors="ignore")
        return {
            "name": name,
            "status": "http_error",
            "status_code": e.code,
            "elapsed": round(elapsed, 3),
            "error": err_body
        }
    except urllib.error.URLError as e:
        elapsed = time.time() - start_time
        return {
            "name": name,
            "status": "url_error",
            "status_code": None,
            "elapsed": round(elapsed, 3),
            "error": str(e.reason)
        }
    except Exception as e:
        elapsed = time.time() - start_time
        return {
            "name": name,
            "status": "timeout_or_unknown",
            "status_code": None,
            "elapsed": round(elapsed, 3),
            "error": str(e)
        }

def generate_markdown_report(results, test_case_name):
    report = []
    report.append(f"# AI Validation & Diff Report: {test_case_name}")
    report.append("\n## 1. Execution Summary\n")
    report.append(
        "| Endpoint | Status | HTTP Code | Latency (s) | Error Details |\n"
        "| :--- | :--- | :--- | :--- | :--- |"
    )

    success_responses = {}

    for r in results:
        name = r["name"]
        status = r["status"]
        code = r["status_code"] if r["status_code"] is not None else "-"
        elapsed = r["elapsed"]
        err = "-"

        if status == "success":
            success_responses[name] = r["response"]
        else:
            err = r.get("error", "Unknown error").replace("\n", " ")

        report.append(f"| {name} | {status} | {code} | {elapsed} | {err} |")

    report.append("\n## 2. Response Outputs\n")
    for r in results:
        report.append(f"### [{r['name']}] Output")
        report.append("```

json")
        report.append(json.dumps(r.get("response") or r.get("error"), ensure_ascii=False, indent=2))
        report.append("

```\n")

    report.append("## 3. Structural & Textual Diff Analysis\n")
    names = list(success_responses.keys())
    if len(names) < 2:
        report.append("Skipping diff analysis because fewer than 2 successful responses were received.\n")
    else:
        for i in range(len(names)):
            for j in range(i + 1, len(names)):
                n1, n2 = names[i], names[j]
                text1 = json.dumps(success_responses[n1], ensure_ascii=False, indent=2).splitlines()
                text2 = json.dumps(success_responses[n2], ensure_ascii=False, indent=2).splitlines()

                diff = list(difflib.unified_diff(text1, text2, fromfile=n1, tofile=n2, lineterm=""))
                report.append(f"### Diff: {n1} vs {n2}")
                if diff:
                    report.append("```

diff")
                    report.extend(diff[:50])
                    if len(diff) > 50:
                        report.append("... (diff truncated)")
                    report.append("

```\n")
                else:
                    report.append("No differences found (Exact match).\n")

    return "\n".join(report)

def main():
    input_data = sys.stdin.read()
    if not input_data.strip():
        print("Error: Test cases must be provided via standard input.", file=sys.stderr)
        sys.exit(1)

    try:
        test_suite = json.loads(input_data)
    except json.JSONDecodeError as e:
        print(f"Error: Failed to parse JSON: {e}", file=sys.stderr)
        sys.exit(1)

    test_name = test_suite.get("test_name", "Unnamed Test")
    endpoints = test_suite.get("endpoints", [])

    if not endpoints:
        print("Error: No target endpoints defined in the test suite.", file=sys.stderr)
        sys.exit(1)

    results = []
    max_workers = min(len(endpoints), 16)

    global_timeout = 10.0
    start_time = time.time()

    with ProcessPoolExecutor(max_workers=max_workers) as executor:
        future_to_endpoint = {executor.submit(call_endpoint, ep): ep for ep in endpoints}

        for future in as_completed(future_to_endpoint):
            ep = future_to_endpoint[future]
            name = ep.get("name", "Unknown")

            elapsed_total = time.time() - start_time
            remaining = max(0.1, global_timeout - elapsed_total)

            try:
                res = future.result(timeout=remaining)
                results.append(res)
            except Exception as e:
                results.append({
                    "name": name,
                    "status": "timeout_or_error",
                    "status_code": None,
                    "elapsed": round(time.time() - start_time, 3),
                    "error": str(e) or "Execution timed out or failed"
                })

    markdown_report = generate_markdown_report(results, test_name)
    print(markdown_report)

if __name__ == "__main__":
    main()
```

💡 **For immediate deployment:** The complete source code suite (ZIP) for this architecture is available on [Gumroad](https://phenox.gumroad.com/l/caicf) for $0+ (Pay What You Want).

After repeating our QA test suite three times, the precise mechanisms that dragged this system into fatal deadlocks and latency traps became evident.

While multiprocessing is highly effective for CPU-intensive tasks, **it is fundamentally unsuitable for network I/O-heavy operations like calling external LLM APIs**. 

The overhead introduced by Inter-Process Communication (IPC), object serialization/deserialization, and OS-level process spawning is not negligible. When multiple external APIs simultaneously experienced delayed responses, the aggressive context switching between processes severely bottlenecked system resources.

To enforce the strict 10-second global limitation, we integrated the following calculation into the core loop:

```
elapsed_total = time.time() - start_time
remaining = max(0.1, global_timeout - elapsed_total)
res = future.result(timeout=remaining)
```

The fatal flaw here is that `as_completed` yields futures *in the order they complete*. 

If the interpreter first evaluates a future from an endpoint that experienced heavy latency (e.g., a local LLM taking 8 seconds to respond), the `remaining` time shrinks drastically. Consequently, when the loop processes the next future—even one from an API that responded almost instantaneously—it applies the severely depleted `remaining` time (e.g., 0.1 seconds). This triggers a **cascading timeout collapse**, forcefully throwing `TimeoutExpired` exceptions and terminating requests that actually succeeded at the network layer.

`asyncio`)
Driven by an absolute constraint to maintain "zero third-party dependencies," we attempted to forcefully parallelize the blocking `urllib.request` using multi-processing, rather than adopting Python's native `asyncio` combined with a non-blocking wrapper (like `http.client` wrapped in an executor) or leveraging asynchronous runtimes. 

By abandoning efficient event-loop multiplexing (which threads or async coroutines handle natively), we constructed a fragile, synchronous wait-state architecture wholly dependent on process scaling—a fundamentally broken approach for scaling HTTP concurrency.

Based on the failures observed in this prototype, we offer the following critical insights for engineers designing similar asynchronous validation pipelines:

`httpx` or native `timeout` parameter in `asyncio.wait()`), ensuring deterministic behavior regardless of the order of completion.
A design that appeared conceptually sound on paper was ultimately shattered against the 10-second barrier by real-world network jitter and a fundamental misapplication of the process concurrency model. However, highlighting the exact boundary limits of standard library concurrency models serves as a solid foundation for more robust architectural decisions in the future.

In software engineering, the structural understanding of *why* failing code breaks is the most valuable asset for future success. We hope this postmortem acts as a navigational compass for developers stepping into the demanding territory of highly concurrent API validation.

*If this engineering log saved your production server (and your sanity), consider supporting our architecture on GitHub Sponsors.*
