No exploit. No malware. No system vulnerability. Just a video call that looked like any other meeting on the calendar.
January 2024 - An employee of the British design and engineering firm Arup – responsible for iconic projects such as the Sydney Opera House, the Bird's Nest Stadium from the 2008 Beijing Olympics, or the Crossrail transport programme in London – receives a phishing email. He's skeptical. Shortly after, a video call follows with several participants, apparently including employees of the company. Everyone looks like colleagues, everyone sounds like colleagues, so he sets his doubts aside.
What he didn't know: every person on that call except himself was an AI-generated deepfake. Deepfake usually refers to fabricated videos in which artificial intelligence is used to make someone appear to do things that never happened, while looking extremely realistic. During a seemingly genuine corporate video conference, the fraudsters posed as senior executives and convinced him to make several transfers to specific bank accounts. A total of 15 transactions, spread over several days – ultimately resulting in a loss of around 25.6 million US dollars.
This incident is not an isolated case. In a similar incident, Mark Read, CEO of the advertising holding company WPP, was already targeted by a very similar scam attempt – using an AI-generated voice clone of himself. Invoice fraud, phishing, WhatsApp voice spoofing, deepfake video calls: companies worldwide are now regularly confronted with this type of attack. For a simple reason: attackers don't need to develop elaborate, complicated exploits – they simply rely on straightforward psychological social engineering tactics.
A victim who falls for the deception acts in good faith, believing they are doing the right thing. In reality, this plays right into the perpetrator's actual motive: harvesting credentials or triggering a wire transfer – often the gateway into an otherwise well-protected corporate network.
Voice phishing is now considered the second most common initial attack vector for gaining access to corporate networks; contact centers alone are estimated to be exposed to a fraud volume of 44.5 billion US dollars in 2025. The voice channel is no longer a fallback channel – it is the front door.
What vishing really is – and how it has evolved into deepfake vishing
Vishing – short for voice phishing – is a subcategory of social engineering. The term refers to fraudulent phone or video calls and voice messages designed to trick a victim into providing sensitive information such as credentials, credit card numbers, bank details, or wire transfers. Attackers exploit the fact that many employees have little time in their daily work routine and readily respond to seemingly trustworthy requests without questioning the caller's legitimacy.
Social engineering in general targets the "human factor" – the supposedly weakest link in any security chain. What gets exploited are fundamentally human traits such as helpfulness, trust, or fear. According to a study by Barracuda (as of June 2021), every organization faces an average of over 700 social engineering attacks per year – the majority of which are phishing (43%) and classic scams (39%), followed by business email compromise (10%) and extortion (2%).
What has fundamentally changed in recent years is the technical quality of the attack tool itself:
Social engineering cannot easily be prevented. No matter how high the technical security measures are – systems cannot be protected if the people who have access to them willingly hand over credentials.
Four phases: How a deepfake vishing attack is built
Unlike a technical exploit, a deepfake vishing attack doesn't need a vulnerability in the system – just preparation, material, and a credible scenario. The actual call is usually the shortest part of the whole process.
The psychological weapon: Cialdini's principles applied to vishing
The technology behind deepfake vishing is new. The psychological levers that carry the actual fraud are not. Robert Cialdini's seven principles of persuasion explain why even attentive, well-trained employees cooperate in the decisive moment — as we’ve already shown using the examples of the MGM hack and modern vishing attacks.
| Principle | Mechanism of Action | Vishing Example |
|---|---|---|
| Authority | People follow authority without questioning it. Attackers pose as supervisors, IT admins, government officials, or auditors. | "This is Thomas Bergmann, CTO. I need your access code for the emergency server right now – our security system has triggered an alert!" |
| Urgency | Time pressure prevents rational thinking. Anyone with no time to think doesn't question. Used in almost every vishing attack. | "The money has to go out within 30 minutes, or the deal falls through." |
| When others have already taken an action, it automatically appears legitimate. | "All your colleagues in the department have already confirmed their data – you're the last one missing." | |
| Liking | People are more likely to help those they like or perceive as similar to themselves. Shared interests are researched in advance on LinkedIn. | "We're colleagues from the same conference, aren't we!" |
| Reciprocity | People feel obligated to return favors. The attacker helps first, then "casually" follows up with the actual request. | Offering help with an IT problem first, then casually asking for credentials. |
| Once someone has said yes, they're more likely to say yes again. First small, harmless requests, then the actual demand. | "That wasn't a problem, was it? Then surely you can also..." | |
| Unity / Belonging | A sense of "we" and group identity is exploited – especially effective in combination with authority. | "As a member of the leadership team, we need your help right now..." |
In vishing attacks, simply invoking an authority role, combined with the correct name of a supervisor – researched on LinkedIn – is often enough to get employees to cooperate. A particularly common variant is IT support impersonation: the attacker poses as a helpdesk employee, cites a real ticket number and department name – and thereby appears fully legitimate.
Why technical filters fail – and how companies can really protect themselves
The core vulnerability isn't in the network, but in the brain: it's wired not to question familiar voices and familiar faces. Effective defense combines processes, technology, and awareness:
- Safewords / code words for critical transactions: a word agreed on in advance that gets requested for transfers or access requests – regardless of how convincing the voice or image seems.
- Voice authentication protocols & detection software: technical systems that can detect synthetic audio artifacts in real time.
- Employee training and simulations: targeted awareness of the manipulation of emergency scenarios. The core feature of every social engineering attack is the deception about the perpetrator's identity and intent – that's exactly what needs to be trained, not just "spotting phishing links."
- Multi-factor authentication (MFA): as an additional barrier, even if credentials or approval were obtained through deception.
- Actively question identities: back-verification through a second or multiple independent channels – not through the number or contact from the suspicious call itself.
- Responsible use of social media profiles: publicly available audio, video, and organizational material is the first gateway for attackers.
For companies with many employees – like Arup – targeted awareness training is particularly effective: those who know what phishing messages, deepfakes, and other forms of social engineering typically look like are more likely to spot the warning signs before the damage occurs.
Incident response: What to do if it has already happened?
- Notify IT security
- Stop using the affected system, or take it offline if possible
- Change passwords on other devices right away
- Check MFA tokens: Were any new devices registered?
- Lock the affected account or enforce MFA
- Invalidate session tokens
- Check email forwarding rules – attackers often silently forward emails
- Review permissions: Was additional access, such as calendar access, granted recently?
- Which phishing email or call preceded the incident?
- Were other employees contacted in a similar way?
- Is this a targeted attack on one person or a mass attack?
Conclusion: When AI turns trust itself into a weapon
The seemingly secure walls of global corporations like Arup crumble when AI exploits human trust as a gateway. Classic scams once sufficed – today, a few seconds of audio material and a deceptively real deepfake video call are enough to perfectly impersonate senior executives. The numbers are alarming: vishing activity is exploding by 442%, while manipulated voice clones and video deception threaten companies worldwide with losses in the millions.
Armed with psychological principles such as authority and urgency, attackers exploit the human factor as the supposedly weakest link in the chain. A single unsuspecting employee, trapped in the deceptive belief of doing the right thing, unknowingly transferred 25.6 million dollars to unscrupulous criminals. Technical firewalls fail when the brain doesn't question familiar voices and familiar faces. The voice channel is thus unstoppably becoming the new front door of cybercrime – whose scale organizations can only limit through vigilance, safewords, and multi-channel verification.
How resilient is your team against vishing & deepfake fraud?
We test your team's awareness of vishing and other social engineering tactics – with realistic scenarios instead of checklists.
Request an initial consultation →