arXiv:2609.00718v1 Announce Type: new Abstract: Many automobile and mobility companies deploy learned driving policies on embedded computers with limited memory and power. Pruning, knowledge distillation, and quantization are the standard methods to reduce the size and the inference cost of these policies. However, these methods are commonly assessed by aggregate numerical scores, and such scores may not reflect the ability of the policy to drive safely when interacting with other road users. In this study, we propose a stage-wise closed-loop evaluation approach to follow a driving policy through a compression pipeline. We formulate the driving task as a partially observable Markov decision process (POMDP) and train a belief-state policy with proximal policy optimization (PPO) in Gym-Duckietown. We then extract the actor, compress it one stage at a time, and evaluate it on five driving curricula. We show that structured pruning is the stage at which the driving capability is first lost. Meanwhile, distillation improves the pruned actor, but the improvement is limited by its rehearsal data. Integer quantization of the improved actor loses some of the curricula that require the vehicle to stop and then resume. Interestingly, the same procedure on the unpruned actor preserves all five curricula. Our study thus provides an empirical analysis aiming to answer the currently active discussions on how to accept a compressed driving policy, so as to achieve a safe and statistically reliable deployment of automated driving functions.
A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies
A new study from arXiv (2609.00718v1) finds that structured pruning is the first stage where driving capability is lost in compressed autonomous driving policies, while distillation improves the pruned actor but is limited by rehearsal data, and integer quantization of the improved actor loses curricula requiring stop-and-resume maneuvers. The researchers propose a stage-wise closed-loop evaluation method, training a belief-state policy with PPO in Gym-Duckietown and compressing it stage by stage, showing that the same quantization procedure on the unpruned actor preserves all five driving curricula.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.