{"slug": "event-driven-ml-pipeline-orchestration-for-manufacturing-an-aws-industry", "title": "Event-Driven ML Pipeline Orchestration for Manufacturing: An AWS Industry Experience", "summary": "An arXiv industry experience report (arXiv:2610.06890v1) documents three years of running an event-driven cloud infrastructure for continuous machine learning training in automotive manufacturing, reporting a 72-78% cost reduction versus always-on GPU infrastructure across 40,000+ production training jobs. The system orchestrates GPU-accelerated training of product-specialized model pairs — a physics prediction model and a reinforcement-learning control policy — across multiple plants using Amazon ECS with EC2 GPU capacity providers, SQS messaging with dead-letter queues, an admission-controlled Lambda dispatcher, a Conductor orchestrator on ECS Fargate, and modular Terraform with multi-account separation. A discrete-event simulation found admission control is necessary because naive dispatch loses 65% of jobs, and that queue-draining matches AWS Step Functions latency while eliminating per-job startup overhead; the simulator and Terraform module skeletons are released as open-source artifacts.", "body_md": "arXiv:2610.06890v1 Announce Type: new \nAbstract: We present an industry experience report on three years of operating an event-driven cloud infrastructure for continuous machine learning training in automotive manufacturing. Our system orchestrates GPU-accelerated training of product-specialized model pairs, a physics prediction model and a reinforcement-learning control policy, across multiple plants, coordinating long-running GPU workloads triggered by manufacturing events. The architecture combines Amazon ECS with EC2 GPU capacity providers, SQS-based messaging with dead-letter queues, and an admission-controlled Lambda dispatcher that enforces cluster concurrency limits. A Conductor orchestrator on ECS Fargate initiates dependency-aware retraining chains on a weekly schedule. The entire infrastructure is codified in modular Terraform with multi-account separation. From 40000+ production training jobs we report a 72-78% cost reduction versus always-on GPU infrastructure. A discrete-event simulation confirms that admission control is necessary (naive dispatch loses 65% of jobs) and that queue-draining matches AWS Step Functions latency while eliminating per-job startup overhead. We provide lessons learned and release the simulator and Terraform module skeletons as open-source artifacts.", "url": "https://wpnews.pro/news/event-driven-ml-pipeline-orchestration-for-manufacturing-an-aws-industry", "canonical_source": "https://arxiv.org/abs/2610.06890", "published_at": "2026-10-07 04:00:00+00:00", "updated_at": "2026-10-07 04:16:45.909274+00:00", "lang": "en", "topics": ["machine-learning", "mlops", "ai-infrastructure", "ai-research"], "entities": ["Amazon Web Services", "Amazon ECS", "Amazon EC2", "Amazon SQS", "AWS Lambda", "AWS Fargate", "AWS Step Functions", "Terraform"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/event-driven-ml-pipeline-orchestration-for-manufacturing-an-aws-industry", "markdown": "https://wpnews.pro/news/event-driven-ml-pipeline-orchestration-for-manufacturing-an-aws-industry.md", "text": "https://wpnews.pro/news/event-driven-ml-pipeline-orchestration-for-manufacturing-an-aws-industry.txt", "jsonld": "https://wpnews.pro/news/event-driven-ml-pipeline-orchestration-for-manufacturing-an-aws-industry.jsonld"}}