cd /news/machine-learning/online-security-learning-in-cooperat… · home topics machine-learning article
[ARTICLE · art-89848] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Online Security Learning in Cooperative Multi-Agent Systems under Hidden Byzantine Attacks

A new arXiv study (2608.06520v1) on online cooperative control of multi-agent systems under Byzantine attacks shows that an attacker observing planned actions induces an exact $(s,a)$-rectangular robust Markov decision process (MDP), while a blind attacker induces an $s$-rectangular model. The authors prove that security regret decomposes into return regret plus a cumulative response gap $D_K$, and that two indistinguishable horizon-one instances force $\Omega(K)$ expected security regret even with zero return regret. They also develop a stage-tied robust estimation-to-decisions learner achieving a regret bound of $\widetilde{\mathcal O}(H^2S\sqrt{AK})+\mathbb E[D_K]$.

read1 min views1 publishedAug 10, 2026

arXiv:2608.06520v1 Announce Type: new Abstract: We study online cooperative control of a multi-agent system under Byzantine attacks. Namely, an unknown, fixed subset of agents are Byzantine comprised and can stealthily overwrite its own coordinates of the team's planned joint action after observing that plan. The learner observes planned actions, public rewards, and public states, but neither the overwrite nor the executed joint action. Our objective is security: to optimize the team performance against the worst overwrites and achieve the optimal security value. We first show that the attacker's information determines the geometry. An attacker that observes the planned action induces an exact $(s,a)$-rectangular robust Markov decision process (MDP) whose rows are convex hulls of overwrite-induced public-outcome laws, whereas a blind attacker induces an $s$-rectangular model. We then identify the information-theoretic limit of security learning, showing that the security regret decomposes exactly into return regret against the response generating the data and a cumulative response gap $D_K$. Two indistinguishable horizon-one instances force $\Omega(K)$ expected security regret while return regret is zero, showing that dependence on $D_K$ is unavoidable. Finally, we develop a stage-tied robust estimation-to-decisions learner and prove a regret bound of $\widetilde{\mathcal O}!\left(H^2S\sqrt{AK}\right)+\mathbb E[D_K]$. Our studies thus provide comprehensive theoretical and algorithmic foundations of reliable multi-agent systems under Byzantine attacks.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/online-security-lear…] indexed:0 read:1min 2026-08-10 ·