As malicious phishing campaigns learned to bypass spam filters, security teams spent years training employees to spot badly worded messages and links pointing to spoofed domains. The security awareness training playbook is now running into a threat it was never built for: deepfake audio and video realistic enough that even a familiar voice or face can no longer be trusted as proof of who's on the other end.
The numbers show how fast this shifted. AI-generated media lures are now used in approximately 6.5% of fraud attempts, up from 0.1% three years ago. In Mandiant's M-Trends 2026, voice phishing ranked as the second most common way attackers gained initial access in 2025, making it the single most common route into cloud environments, ahead of email phishing.
When the attack arrives as a voice #
Indeed, as generative AI has gone mainstream, so have use cases that weaponize it to breach corporate systems. In one example from this past winter, a Swiss CEO received several phone calls over the course of weeks, with synthetic audio convincing him to transfer millions to scammers’ bank accounts.
Usually, security awareness training teaches employees to verify suspicious transfer requests using another communication channel, and in this case, the deepfake phone call would have only confirmed it.
Attackers no longer need much raw material to pull this off. Roughly three seconds of audio is often enough to clone a voice convincingly. Traditional phishing training, built around scrutinizing text and links, has little to say about a threat that arrives as audio or video.
What makes this shift harder for cybersecurity teams is speed. Voice cloning and video synthesis tools that once required real production skill are now cheap and accessible.
That gap is forcing a rethink of what security awareness training for employees actually needs to cover. Programs built around static, once-a-year modules are structurally mismatched to a threat that evolves month to month. And yet many companies still have no formal deepfake response plan, leaving security teams looking for training that updates as fast as the attacks do.
Training that updates as fast as the attacks #
The shift also changes what "engagement" means for these programs. A once-a-year compliance module that employees click through to satisfy an audit was already a weak defense against conventional phishing. Against a deepfake voice or video call, it's closer to no defense at all, since the instinct being trained - read carefully, check the sender, hover over the link - doesn't transfer to a spoken request from a face you recognize.
AI-generated simulations, sent to employees over the course of their ongoing work, can go a long way, helping people to recognize patterns and remain vigilant.
Effective programs also need to teach broader skepticism and behavioral change, encouraging employees to verify unusual requests, slow down urgent asks, and treat a familiar voice or face as necessary but not sufficient on its own.
That verification habit is also where policy has to catch up with technology. Some organizations are formalizing a rule that any request involving money movement or credential access gets confirmed through a pre-agreed second channel or verification method that the attacker didn't initiate before anyone acts.
It's a low-tech fix to a high-tech problem, and it works precisely because it doesn't depend on an employee successfully spotting a deepfake in the moment, which even trained people struggle to do consistently.
The organizations that adapt fastest tend to treat awareness training as a continuous, evolving program rather than a fixed curriculum. As deepfake-enabled social engineering becomes routine rather than a novelty, that continuous approach is shaping up to be the dividing line between security teams that stay ahead of the shift and those still training for last year's threats.