You're among CyopScape's first visitors — share your feedback and help us improve.


CyopScape | Cybersecurity Insights Threat Analysis
← Back to Insights
Threat Analysis 9 min read

Voice and Video Aren't Proof Anymore

Understanding the rise of deepfake-enabled social engineering, and why "let's get on a call" is no longer the verification it used to be.

Human verification has always rested on a simple assumption: a voice on the phone, or a face on a video call, belongs to the person it appears to be. That assumption no longer holds. Generative AI has collapsed the cost, time, and skill required to convincingly clone a voice or a face, and attackers are folding that capability directly into social engineering campaigns that once relied on text alone.

The clearest illustration is the January 2024 fraud against engineering firm Arup. An employee in the firm's Hong Kong office was invited onto a video conference with what appeared to be the company's UK-based CFO and several colleagues, all of whom were AI-generated deepfakes built from publicly available meeting footage. Over 15 transactions, the employee transferred roughly HK$200 million (about US$25 million) before discovering the deception by contacting head office.

Earlier voice phishing (vishing) attacks depended on a scammer's own acting ability and a plausible pretext. Deepfake-enabled social engineering replaces the actor with a synthetic reproduction of someone the victim already trusts: a CFO, a colleague, an IT helpdesk technician, even a family member. The result is a new class of attack that defeats the one verification method organizations have leaned on whenever a text-based request feels suspicious: "let's get on a call." In the Arup case, the video call itself is what overrode the employee's initial, correct suspicion that the opening email was a phishing attempt.

This piece looks at how deepfake-enabled social engineering works, why it undermines assumptions baked into most awareness training, and what defenders can do about it.

How the Attack Works

Traditional social engineering follows a recognizable arc: reconnaissance, pretext development, and delivery. Deepfake-enabled attacks follow the same arc, but the pretext stage now includes building a synthetic likeness of a trusted person.

Voice cloning tools can produce a usable synthetic voice from as little as a few seconds to roughly a minute or two of source audio. That material is readily available from earnings calls, webinars, podcasts, town halls, or social media video. Face-cloning and real-time face-swap tools follow a similar pattern, drawing on publicly available headshots, video interviews, or recorded meetings. Arup's attackers built their deepfakes entirely from existing footage of the executives they impersonated.

Once a model is trained on the target's likeness, an attacker can generate synthetic speech on demand, or drive a synthetic face in real time during a live video call, and use it across several delivery channels:

Diagram showing how public voice and video samples train a cloning model that generates synthetic media, delivered through a live call, voicemail, video meeting, or camera injection to pressure a victim into acting
Deepfake-enabled social engineering: from public media samples to a victim acting under pressure

Because the synthetic voice or face is layered onto a live, interactive conversation, the attack inherits everything that makes vishing effective: urgency, authority, and real-time pressure. On top of that, it adds a visual or auditory cue that most people have been trained to treat as strong evidence of identity. Security researchers have documented that malicious actors are now using convincing real-time deepfakes specifically to run phishing attacks during video meetings, not just in pre-recorded clips.

Why This Matters

Security teams have historically treated voice and video as a stronger verification signal than text, precisely because text is easy to fake and a live call "sounds like" the person it claims to be. That assumption is now the vulnerability.

The scale is no longer theoretical. CrowdStrike reported that the use of AI-based voice cloning surged 442% between the first and second halves of 2024. In May 2025, the FBI's Internet Crime Complaint Center issued a public warning about an ongoing campaign using AI-generated voice messages to impersonate senior U.S. officials and their contacts. And on the verification side, threat-intelligence firm Group-IB documented 8,065 biometric injection-attack attempts against a single financial institution over eight months in 2025, all using deepfake imagery fed through virtual cameras.

None of this requires nation-state resources. Consumer-grade voice cloning and face-swap tools are inexpensive, widely available, and require little technical skill to operate. Arup's own CIO later said he was able to build a rough real-time deepfake of himself in about 45 minutes using free, open-source tools. That floor keeps dropping as the tooling improves.

The Broader Security Trend

Deepfake-enabled social engineering extends a pattern seen across recent attacker tradecraft: abusing something a target already trusts rather than building new infrastructure to attack it directly. Just as attackers have learned to route traffic through trusted cloud platforms or messaging services to blend in with normal activity, they are now borrowing the identity of trusted people to blend in with normal conversation. As one analysis of the Arup incident put it, the controls organizations have invested in most heavily, such as MFA, EDR, email filters, and network monitoring, are largely the wrong shape for an attack that targets a decision rather than a system.

This trend is accelerating alongside two developments. First, real-time voice and video conversion has moved from research demonstrations to consumer-accessible tools, lowering the bar for live impersonation during calls and meetings. Second, organizations continue to expand the use of voice and video as identity verification for high-value actions, often as a reaction to distrust in email and text, without accounting for the fact that the newer channel carries its own, less familiar forgery risk. Injection attacks that target this exact gap are rising fast: iProov's 2025 threat reporting recorded a 2,665% surge in native virtual-camera attacks against verification systems.

As these tools mature, defenders should expect deepfake components to appear as one stage within a larger multi-channel campaign rather than as a stand-alone attack. A deepfake might follow an initial phishing or smishing message, or precede a request to install a tool or grant access. The FBI's advisory describes exactly this pattern: a smishing text to establish rapport, followed by an AI-generated voice message, followed by a pivot to another platform where credentials are harvested.

Defensive Takeaways

Retire "voice or video equals proof"

Treat a phone call or video meeting as one input among several, not as sufficient verification on its own for high-risk requests.

Require out-of-band verification for sensitive actions

Wire transfers, credential resets, access grants, and vendor payment changes should require confirmation through a second, independently established channel, not a number or contact provided during the same interaction. In the Arup case, a single verified callback would have exposed the fraud.

Establish pre-agreed callback procedures

Verify unexpected high-risk requests by calling back on a number sourced from an internal directory, not one provided by the caller or in the suspicious message.

Use codewords for critical approvals

The FBI recommends pre-shared verbal codewords to verify identity in voice messages, a check that publicly available audio cannot reproduce. Extend the same idea to executive-approval workflows.

Update awareness training to cover voice and video

Extend phishing and vishing training to include deepfake indicators, and run simulations that include a live or recorded voice or video component rather than email alone.

Harden identity-verification flows against injection

Where video is used for onboarding or KYC, deploy liveness and camera-integrity checks that detect virtual cameras and injected streams, not just presentation attacks held up to a lens.

Log and review high-risk approval requests

Monitor patterns in wire-transfer approvals, credential-reset requests, and access changes for unusual urgency, off-hours timing, or requests that bypass normal channels.

Final Thoughts

The security industry spent the last decade teaching people to distrust text: suspicious links, spoofed domains, and poorly worded emails. Deepfake-enabled social engineering targets the channel that training never covered: the sound of a familiar voice, or the sight of a familiar face on a screen. As Arup's CIO framed it, the assumption that "I spoke to them" equals confirmation is now expiring, and not in some distant future but in current incident reports.

This does not mean voice and video verification are worthless. It means they can no longer be treated as sufficient on their own for high-stakes decisions. As with other AI-enabled threats, the most durable defense is procedural rather than perceptual: build verification steps that do not depend on trusting the channel in the moment, and treat urgency itself as a signal worth slowing down for.

← Back to Insights