A few seconds of voice recording. A single photo. That's all it takes today to clone a voice or a face convincingly.
May 2024: Mark Read, CEO of the advertising holding company WPP, becomes the target of an AI-generated voice clone of himself – fraudsters attempt to deceive colleagues using his cloned voice and a fake video call.
January 2024: At the British design and engineering firm Arup, a highly sophisticated video vishing attack leads an employee to transfer funds after being instructed to do so during a fake video conference with supposed executives.
June 2026: At the IT service provider MSG, a vishing attack results in a data leak of 26 million records.
Three completely different companies, three completely different industries – one common denominator: deepfake technology, used as a tool for social engineering. Vishing has therefore long ceased to be a simple phone call and has become, in the age of generative AI, one of the most effective attack surfaces of all – and the numbers confirm just how fast this is developing.
These figures raise two central questions that we explore in this article, both practically and technically: how low is the barrier to entry for attackers really – and how do deepfakes work under the hood? To answer this, we didn't cite a statistic – we created a deepfake ourselves.
What exactly are deepfakes?
The term is a combination of Deep Learning and Fake. It refers to generated video, image, or audio material that uses AI to manipulate human faces or voices so that one person is replaced by another or convincingly simulated – visually and acoustically almost indistinguishable from the original.
The technology itself is neutral. It's the application that determines whether it results in benefit or harm:
The same technology that gives a person their own voice back can, in the wrong hands, be used to impersonate someone else. The difference lies not in the tool, but in the intent behind it.
Image, audio, or video – what kind of deepfake should it be?
Before a deepfake is created, a simple decision has to be made: which medium should be manipulated? The three common categories differ significantly in effort and the source material required.
Tried it ourselves: how easy is a deepfake really?
To honestly assess the barrier to entry, we didn't cite a report – we got our hands dirty ourselves. Goal: a short video in which the British actor Tom Holland – known for his lead role as Spider-Man – says something he never said. The platform ElevenLabs was used for this.
The generated script for the demo clip was deliberately worded so that the fake reveals itself at the end:
The entire process – from voice recording to the finished, talking video – could be carried out using freely accessible tools and without specialist technical knowledge. That's the real takeaway: the barrier today no longer lies in the technology, but at most in the price.
How a classic deepfake is technically created
Behind the simple user interface lies a multi-stage process built on computer vision and neural networks. Using classic face-swapping as an example, this can be broken down into four steps.
What attackers actually use deepfakes for today
These technical capabilities translate directly into concrete attack scenarios – from bypassing biometric security systems to targeted reputational damage.
| Domain | Attack Objective | Description |
|---|---|---|
| Authentication | Defeating Biometric Systems | Since fake voices and faces can now sometimes be generated live, deepfakes effortlessly bypass digital security controls such as telephone voice recognition – the system simply lacks any way to look behind the curtain remotely. |
| Targeted Fraud | Using deceptively realistic voices and images, fraudsters trick victims by, for example, perfectly imitating their boss on the phone and thereby forcing blind obedience for costly wire transfers. | |
| Public Discourse | Disinformation | Fake videos of politicians or celebrities can create a perceived truth online within seconds, deliberately deceiving and manipulating entire societies. |
| Reputation | Defamation | Because deepfakes can attribute any scandal or statement to any person, a single fake video is enough to permanently ruin a victim's reputation. |
How to spot deepfakes – what to actually look out for
Awareness is the key countermeasure: anyone who knows how such an attack works technically can assess the authenticity of image, video, and audio material far more accurately. Even today, most deepfakes still leave behind typical, recurring artifacts.
No single characteristic on its own proves a fake. Only the combination of several small inconsistencies – both visual and acoustic – produces a reliable picture. For security-critical decisions, technical material should therefore never be the sole basis for verification.
Conclusion: the barrier has fallen – vigilance must rise
Our own test makes it clear: a convincing deepfake today is no longer a matter of elaborate specialist software, but a matter of a few seconds of source material and a freely accessible tool. What was once reserved for Hollywood studios is now within reach of a few clicks – and that's exactly what makes vishing and other deepfake-powered attacks so dangerously effective.
Technical understanding remains the best protection: anyone who knows how encoders, decoders, and lip-sync models work together also recognizes the limits of the technology – the subtle artifacts that still give fakes away.
Further reading on this topic:
- Deepfake Vishing – how vishing works in practice through digital channels
- Pretexting – how cover stories are built for physical attacks
- MGM Hack – what vishing can do on a large scale
Would your team recognize a deepfake call?
We use realistic vishing and deepfake scenarios to test whether your employees and processes hold up in a real incident – and show you in the report exactly where the gap lies.
Request an initial consultation →