In January 2024, an employee at the finance department of engineering firm Arup’s Hong Kong office joined what looked like a routine video call with the company’s UK-based chief financial officer and several other colleagues. Every other person on the call was a deepfake, generated from publicly available video and audio of the real executives. Believing the instructions were genuine, the employee authorized 15 transactions totaling roughly HK$200 million, about $25 million, to accounts controlled by the attackers, according to Hong Kong police and Arup’s own confirmation to CNN.

Arup confirmed the incident publicly in May 2024, saying it had reported the fraud to Hong Kong police and that none of its systems were compromised — the attack relied entirely on synthetic video and audio, not a network intrusion. It remains one of the largest publicly confirmed corporate losses to date attributable to generative AI impersonation used in a live, interactive video call rather than a pre-recorded clip or a single spoofed phone call.

The attempt that failed: WPP

Not every attempt succeeds, and the near-miss is instructive. In May 2024, scammers targeted advertising group WPP by creating a WhatsApp account with a photo of chief executive Mark Read, setting up a Microsoft Teams meeting, and using AI-generated voice cloned from YouTube footage of Read plus a deepfaked video image to try to convince an agency leader to set up a new business and transfer funds, according to reporting at the time. The scam failed because the targeted employee grew suspicious during the call and contacted Read through a separate, verified channel before acting. WPP confirmed the attempt publicly; no funds were lost.

Why the video-call format works

Both cases follow the same mechanics that make deepfake fraud different from older forms of business email compromise (BEC). Instead of a spoofed email domain or a single phone call, attackers stage a live or apparently live video meeting using a handful of real public video clips — earnings calls, conference talks, YouTube interviews — run through voice-cloning and face-swapping tools. The live, multi-participant format lends the request social proof that a text message or single voice call does not carry, which is exactly why the Arup employee proceeded despite what employees later described as some hesitation about the unusual payment request.

Controls that have actually stopped these attacks

What separates the WPP near-miss from the Arup loss is not detection technology — it’s a verification step outside the channel the attacker controls. Security researchers and the incident writeups from both cases point to the same small set of measures:

  • Out-of-band verification for any payment instruction above a set threshold: calling back a known, previously verified phone number, not one supplied in the meeting or email.
  • Callback and dual-approval rules that require a second, independently contacted approver for wire transfers over a defined amount, regardless of how senior the requester on the call appears to be.
  • Pre-agreed verification phrases or challenge questions for high-value approvals, which are difficult for a real-time deepfake to answer correctly without preparation.
  • Employee awareness training that explicitly covers video-call impersonation, not just email phishing — most corporate training as of 2024 still treated deepfakes as a hypothetical rather than a live threat vector.

The trend since Arup

Since the Arup case became public, security vendors and law firms tracking incidents — including entries in the AI Incident Database and the OECD’s AI incident monitor — have logged a steady stream of similar attempts against finance and HR staff, ranging from fabricated job-candidate interviews to CEO voice-clone calls requesting urgent wires. The common thread across the confirmed cases is not the sophistication of the deepfake itself but the absence of a verification step that does not depend on trusting what the caller says or how they look on screen.

What makes the Arup and WPP cases useful as reference points, rather than just alarming headlines, is that both are independently confirmed by the targeted companies themselves rather than relying on anonymous or third-party accounts. Arup told CNN directly that it had reported the matter to Hong Kong police and that the loss was real; WPP’s chief executive addressed the attempted scam publicly rather than letting it circulate as rumor. That level of corporate confirmation remains rare in this category of fraud, since most targeted companies have no legal obligation to disclose an attempted or even successful deepfake scam unless it also triggers securities-disclosure or data-breach reporting rules, and reputational concerns push many toward silence.

The economics that make this attractive to criminals

The tooling required to stage a convincing video-call impersonation has become markedly cheaper and more accessible since the Arup incident. Commercial and open-source voice-cloning models can now produce a usable clone from a few minutes of source audio — readily available for any public company executive who has given earnings calls, keynote talks or media interviews — and real-time face-swapping software can run during a live video call rather than requiring pre-rendered footage. That combination lowers the cost of an attempt to roughly the price of research and setup time, while the potential payout, as Arup demonstrated, can run into eight figures from a single successful call. Security researchers tracking the category generally agree the return on investment for attackers is why deepfake-enabled fraud attempts against finance and HR functions have continued to appear through 2025 and into 2026, even as awareness of the Arup case has spread widely within corporate security teams.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *