A deepfake is an image, video, or voice recording generated or altered by a neural network so convincingly that a viewer or listener mistakes it for an authentic recording of a real person. The technology behind it is generative AI: an algorithm trains on real photos, video, or audio of a person, then reproduces that person’s face, voice, or manner of speaking in a scene that never actually happened.
What Is a Deepfake in Simple Terms
The word “deepfake” is a blend of two terms: deep learning and fake. It emerged in 2017 on the Reddit forum, where a user going by the name deepfakes posted videos with swapped faces — and since then the word has stuck as the name for this entire class of synthetic media technology.
A deepfake is not the same thing as an ordinary Photoshop edit or video splice. The difference is who does the work. A Photoshopped forgery is assembled by hand, frame by frame, and the more elaborate the forgery, the more time and skill it takes. A deepfake is produced by a neural network: it trains itself on a set of images or voice recordings of a specific person and then generates new content — as realistic as its training allows. The barrier to entry for fraudsters has dropped along with the cost: BI.Zone specialists estimate that producing one deepfake video takes 5–7 minutes of work at a cost of roughly 50 rubles.
The technology itself is neutral — it is used in filmmaking, dubbing, advertising, and restoring the voices of people who have lost speech due to illness. The problem is that the same tool falls just as easily into the hands of fraudsters.
How Deepfakes Are Made and Why It Became Widespread
Different types of deepfakes are built on different neural network architectures. The first generation of these models was usually built on GANs — generative adversarial networks, where one network (the “forger”) generates the fake and a second (the “detective”) tries to spot it; after every failed detection attempt, the forger learns to do better, and the quality of the fake gradually improves. Diffusion models work differently: without a forger-detective pair, they reconstruct an image, video, or sound from noise through a long sequence of steps — and in practice this often produces a higher-quality, more controllable result than a classic GAN. Readers don’t need to dig into the math behind these models — what matters is that, over the past few years, tools built on both architectures have multiplied many times over, and by default they require no special skill to use.
In practice, a deepfake attack relies on one of several methods, and often combines them:
Replacing one person’s face with another’s in a video or photo. The most recognizable vector: according to observations by the iProov research center (iSOC) within the stream of attacks on remote identity verification, the number of face-swap attempts grew by 300% in 2024 compared with 2023.
Transferring one face’s expressions and movements onto another. Used to “animate” a static image or make someone else’s face appear to say the words the attacker wants.
Synthesizing a specific person’s voice. A recognizable voice copy can be produced from just a few seconds of clean recording: various estimates put this at 3 to 15 seconds, though the exact threshold depends on the quality of the source recording, the language, the specific model, and the required degree of similarity.
Swapping a face or voice live, during a video call or conversation, with no pre-recording involved. This is the method behind the highest-profile corporate cases of recent years.
Virtual camera injection deserves separate mention — a method of feeding a pre-prepared or generated video stream into a system so that it appears to come from the device’s real camera. Here the quality of the deepfake itself is secondary: the core problem is that the forgery bypasses the optical path entirely — the system receives a ready-made file instead of a shot of a live person. According to the same iProov observations, in 2024 versus 2023 the number of virtual camera injections within the stream of attacks on remote verification grew by 2,665% — the fastest-growing vector in that specific dataset; this percentage cannot be extrapolated to the cybercrime market as a whole. A detailed breakdown of the technical layers that address this specific threat is available in the article “How Video Streams Are Spoofed in KYC”.
Types of Deepfakes
Below are the types of synthetic content — that is, what exactly the neural network generates:
| Type | What It Does | Where It’s Most Often Used Against Businesses |
|---|---|---|
| Face swap | Replaces a person’s face in a video or photo | Video calls, fabricated “evidence” |
| Face reenactment | Transfers expressions and speech onto another face | Synthetic statements, compromising material |
| Voice cloning | Synthesizes a voice from a short sample | Fraudulent calls, voice messages from a “manager” |
| Real-time deepfake | Swaps a face/voice live, during a call | Video conferences, calls with “colleagues” and “management” |
Separate from the content itself is the method used to present it to a verification system — virtual camera injection, a presentation attack, or a replay attack (more on this below, in the KYC section) — and the fraud scenario the deepfake is embedded in: for example, synthetic identity, where a fraudster combines other people’s data with fabricated data and deepfake biometrics to open fraudulent accounts or take out loans.
Why Deepfakes Are Dangerous for Business
A deepfake rarely exists in isolation — the threat begins where the technology meets ordinary human trust. Fraudsters don’t break a company’s technical defenses; they deceive a specific employee who makes a decision under the pressure of urgency and authority.
For a business, this translates into several kinds of risk at once:
- Financial. A direct money transfer made on instructions from a deepfake copy of an executive — the employee is convinced they’re following a real person’s order. Further in the article — real cases with specific amounts.
- Reputational. A synthetic video or audio clip of a company executive making a “statement” they never actually made — the reputational hit lands faster than the company can debunk the forgery.
- Operational. A deepfake passes verification when opening an account or accessing a service — and a fraudulent account appears in the system, which a real criminal then goes on to use.
- Compliance. Fraud missed during onboarding calls the AML/KYC procedures themselves into question before the regulator — a risk broader than the sum of specific losses.
- HR and staffing. A deepfake interview — when a synthetic face and voice show up for a call with a recruiter. In this scheme, the attacker is usually after access to the company’s internal network through the new hire’s position — the salary is secondary.
A separate category is attacks via video calls and social engineering. Here the deepfake works as a psychological lever: a familiar voice and face belonging to an executive sharply lower an employee’s guard, and the technical sophistication of the forgery becomes secondary. This is precisely the vector behind the largest confirmed losses in recent years.
We’ll help you identify which layers of deepfake protection your business actually needs — from liveness checks to payment confirmation procedures
Real Fraud Scenarios
Deepfake fraud has stopped being a theoretical risk — below are cases confirmed by independent media and industry reports: specific investigations naming companies and exact damage amounts.
Hong Kong, engineering firm Arup, 2024. An employee in the finance department was invited to a video call attended by several “colleagues,” including the CFO — all of them turned out to be deepfake reconstructions of real employees, built from their public appearances. Convinced he was seeing and hearing familiar colleagues, the employee transferred $25.6 million to the fraudsters’ accounts (CNN, CrowdStrike Global Threat Report 2025). Takeaway for business: a video call on a familiar platform no longer proves, by itself, who you’re talking to — payment authorization must not rely solely on “I saw and heard my boss.”
Germany/UK, energy company, 2019. One of the first documented cases of voice-clone fraud: an employee at the UK subsidiary received a call from the “voice” of the German parent company’s CEO, demanding an urgent transfer of €220,000 to a “Hungarian supplier’s” account (The Wall Street Journal). The voice was cloned so precisely that the employee had no doubts about the accent or intonation. Takeaway: even a single, maximally convincing call must not serve as the sole basis for a payment.
Ferrari, 2024. An employee received a call purporting to be from CEO Benedetto Vigna — the voice was convincingly faked, but the employee asked a follow-up question about a book the executive had recently mentioned, and the fraudster had no answer (Fortune). The attack was stopped by exactly this simple procedural trick. Takeaway: a pre-agreed, “off-script” verification method is an effective barrier even against a high-quality voice clone.
The same logic worked in the case of advertising holding company WPP that same year, 2024: an attempted deepfake attack via a Microsoft Teams video call using a cloned executive voice was detected and stopped by an additional verification procedure — not by an employee’s visual vigilance (OECD.AI).
How Deepfakes Are Used Against KYC and Identity Verification
Remote identity verification (onboarding, video KYC) is usually built on a chain: the user films themselves on the device’s camera, and the system matches the face against the photo in the document and checks for signs of a live presence. A deepfake attacks precisely this chain — and not in just one way.
A simple comparison of the face against the document photo has long been insufficient — the attacker presents the camera with a synthetic or someone else’s face, fitted to the required document, instead of their own. Here it’s important to distinguish the specific method by which the deepfake content is presented to the system:
- Presentation attack — a photo, a video on another device’s screen, a mask, or a deepfake video played back on a separate screen is physically shown to the camera instead of a live person.
- Replay attack — a special case of a presentation or injection scenario in which a previously intercepted legitimate recording is fed back into the system; the specific class depends on whether the recording is shown to the camera physically (presentation) or fed in programmatically, bypassing the camera (injection).
- Virtual camera injection — a ready-made video stream, including real-time face swap, is fed directly into the application by software means, bypassing the device’s physical camera; to the system, this looks like an ordinary signal from the sensor.
Separate from the presentation method is the synthetic identity fraud scenario — where an attacker combines fragmentary real data about a person (or entirely fabricated data) with deepfake biometrics to create a profile that doesn’t correspond to any real person but appears “verifiable,” for long-term use in fraudulent operations.
An Entrust report published in 2026, analyzing 2025 data, shows that within its dataset, deepfakes account for roughly one in five biometric fraud attempts. Voice identification is subject to a similar problem: on its own, without additional protection, basic voice verification does a poor job of telling a cloned voice apart from a real one — more on this below, in the section on voice deepfakes.
A detailed, frame-by-frame breakdown of how face swapping is detected in a KYC video stream — by motion dynamics and codec artifacts — is available in the article “Deepfakes in KYC: Detecting Face Spoofing and Protecting Video Verification”. And what actually happens at the SDK and network-channel level during a synthetic video injection is covered in “How Video Streams Are Spoofed in KYC”.
Voice Deepfakes: A Closer Look
The voice channel shows one of the sharpest spikes: according to Pindrop, within the analyzed set of calls, attempts at voice fraud using deepfakes grew by more than 1,300% in 2024. Voice cloning has become accessible precisely because of its low barrier to entry: getting a working copy doesn’t require access to expensive studio recordings — a short clip from an interview, a webinar, a story, or a voice message the person left publicly themselves is enough.
The mechanics of voice fraud usually follow the “fake boss” pattern: an employee receives a voice message or call from a “manager” demanding an urgent transfer of money to new account details, with no discussion with anyone else. A familiar voice sharply raises trust in the request — in a situation where a text message from an unfamiliar address would raise suspicion immediately.
The problem is that human hearing does a poor job of recognizing a high-quality voice clone — intonation, pauses, and even breathing patterns are reproduced accurately enough to fool someone who knows the original voice well. Voice ID without additional protection is vulnerable to this kind of forgery: technically, this vulnerability is reduced by voice anti-spoofing and voice liveness (verifying a live voice rather than a played-back recording), as well as challenge-response — asking the speaker to say a random phrase or number on the spot — and analysis of the call’s risk signals. Separate from technical protection, a procedural safeguard also works: a callback to a number known in advance, a code word that no outsider could know, and a rule requiring a second person to confirm the payment.
Can You Spot a Deepfake on Your Own
The question of how to spot a deepfake on your own deserves separate treatment, because a widespread misconception has built up around it. For a long time, advice circulated online: watch for blinking, seams around the face, odd lighting, or lips out of sync with sound. The problem is that these are signs of earlier generations of deepfakes — modern generative models and post-processing are noticeably better at removing the artifacts that used to give a forgery away.
Detection specialists at Metaphysic warn directly: tricks like asking someone to turn their head in profile or wave a hand in front of their face still work for now, but neural networks are improving quickly, and such checks are losing their ability to catch a forgery. The once-popular “three-finger test” is already unreliable and gives people a false sense of security.
A meta-analysis by Diel et al. (Computers in Human Behavior Reports, 2024) showed just how poorly people perform at spotting deepfakes: 56 academic publications and 137 individual effect sizes across several modalities — not only visual recognition by eye, but also perception of audio and combined content. The average human accuracy at identifying a deepfake across this dataset was 55.5%, and the result’s confidence interval crosses the 50% mark — meaning statistically it’s close to plain guessing. A separate iProov study examined recognition on material that included both images and video — and only 0.1% of participants correctly classified the entire set of material presented to them.
Visual and auditory checks by a human are, at best, a weak supplementary layer of protection. The main load must be carried by systematic verification through process and technology, covered below — personal attention to the details of a video guarantees almost nothing here.
How to Protect Against Deepfakes: Layered Defense for Business
The core principle of deepfake protection is layering. No single method covers the entire range of threats, and presenting one technology as a complete solution is a common and dangerous mistake.
| Method | What It Protects Against | What It Doesn’t Solve on Its Own |
|---|---|---|
| Liveness detection | Presentation attacks: photos, video played on a screen, masks | Doesn’t always catch virtual camera injection without a separate detector |
| Injection attack detection | Video stream spoofing via a virtual camera, bypassing the physical camera | Doesn’t replace checking whether the face itself matches the document |
| Anti-spoofing (ISO/IEC 30107-3 testing standard, PAD testing at the independent lab iBeta) | Formalizes and confirms resilience to presentation attacks | Doesn’t cover voice attacks or attacks via messaging apps |
| Document verification | Forged or edited identity documents | Doesn’t detect a deepfake if the document is genuine but the person is not who they claim to be |
| Biometric verification (1:1) | A face that doesn’t match the document photo | Vulnerable to a high-quality face swap without a separate anti-spoofing layer |
| Device intelligence | Emulators, suspicious device configurations, root/jailbreak | Doesn’t analyze the content of the video or audio itself |
| Behavioral analysis and risk scoring | Anomalous action patterns that raise the likelihood of fraud | Requires data to build up a profile; doesn’t work instantly on a new user |
| MFA and re-authentication | Compromise of a single confirmation channel | Doesn’t help if several factors are compromised at once |
| Verification via an independent channel (callback, code word) | Voice and video calls from a “manager” or “employee” | Requires an established procedure and the discipline to follow it |
| Transaction monitoring (anti-fraud loop) | Suspicious operations after a successful attack on verification | Doesn’t prevent the attack itself at the identification stage |
| Employee training | Reduces the likelihood that an employee will give in to pressure created by urgency | Doesn’t protect against technically sophisticated attacks with no obvious signs |
For payments and instructions issued in the name of management specifically, a procedural rule applies: confirmation via an independent communication channel the attacker has no access to. A callback to a number from the corporate directory (from the directory — not the number the call came from), a code word known to a narrow circle of people, and mandatory sign-off on the payment by a second employee are inexpensive measures that stop a significant share of “fake boss” attacks before the money ever leaves.
How Liveness and Anti-Spoofing Help During Verification
Liveness detection verifies the presence of a live person in front of the camera at the moment of capture — as opposed to a photograph, a video played on a screen, a mask, or a synthetic forgery. There are two main modes: active liveness, which asks the user to perform an action — turn their head, say a random number — and passive liveness, which analyzes a single frame or a short sequence with no extra action required from the person.
It’s on the liveness side that a formal methodology for assessing resilience to forgery has been built. The international standard ISO/IEC 30107-3 sets the evaluation criteria for PAD (Presentation Attack Detection) — a common set of metrics for how resilient a system is to different categories of attack. Independent testing laboratories, such as iBeta, evaluate such systems using their own PAD testing methodology.
Liveness has clear limits: it does a good job of covering presentation attacks — photos, videos, and masks shown to the camera — and it remains one of several layers of protection. Voice cloning, attacks via messaging apps, or fully synthetic video calls without direct video verification require the other layers of protection listed in the table above. A breakdown of the specific technical methods attackers use to bypass liveness, and why it doesn’t always work, is available in the article “How Liveness Detection Is Bypassed”.
NeuroVision’s liveness check runs in passive mode, requiring no extra action from the user, has been tested for resilience to spoofing attacks against the requirements of the ISO/IEC 30107-3 standard, ships as a REST API and SDK for mobile and web applications, and deploys in the cloud or on-premises — detecting masks, photos, videos, and spoofing attacks, including deepfakes, within a single protective loop alongside the other verification layers. NeuroVision’s face recognition algorithms (under the Enface brand) have been tested in the international NIST FRVT benchmark.
NeuroVision’s liveness check detects masks, photos, videos, and deepfakes in a second — in passive mode, with no extra action required from the user
Law and Regulation
In the European Union, Article 50 of the AI regulation (the EU AI Act, Regulation (EU) 2024/1689) has been in effect since 2 August 2026, and it separates two distinct obligations for different market participants. Article 50(4) requires persons and organizations that deploy such systems (deployers) to disclose to users that specific content is a deepfake, if it is presented as a genuine image, video, or audio recording. Article 50(2) separately requires providers of synthetic content generation systems to mark that content with machine-readable labeling that allows its artificial origin to be identified. The regulation provides an exemption for use permitted by law for law-enforcement purposes, as well as a special, lighter-touch disclosure regime for artistic, creative, satirical works and similar cases.
Russia currently has no dedicated law specifically on deepfakes. In August 2026, a bill was submitted to the State Duma proposing amendments to Article 63 of the Criminal Code of the Russian Federation — the list of circumstances that aggravate criminal liability. The proposed logic is to add to that list the use of AI to create material that imitates a specific person’s appearance, voice, or statements, as well as the creation of images of non-existent people: this would function as an aggravating circumstance for existing offenses, without introducing a separate new offense specifically “for deepfakes.” As of the time this material was prepared, this is a bill, not an enacted law, and its wording may change during review.
Separately, Federal Law No. 210-FZ — the “second anti-fraud package” — is in force. It was passed by the State Duma in June 2026 and expands measures against cybercrime. An important nuance: the key new provisions of this law, including a mechanism for reimbursing stolen funds and expanded requirements for banks on transfers made without the customer’s consent, take effect only from 1 March 2027 — as of the time this material was prepared, 21 August 2026, these provisions are not yet in force. The regulation itself is broader than the topic of deepfakes and covers cybercrime as a whole, including cases involving synthetic content.
In addition, Russia has Federal Law No. 152-FZ “On Personal Data” in force: a facial photo or a voice recording is treated as biometric personal data specifically when the operator uses it to establish a person’s identity — not in any context by default. A leak of a database of customers’ voice or video recordings collected specifically for identification falls under the requirements and penalties of this law — regardless of whether that data was subsequently used to make deepfakes.
Business Protection Checklist
- Implement a callback to a known number — to confirm any non-standard payment order received by voice, video, or messaging app.
- Set up a code word — for critical instructions, shared within a narrow circle of employees, and never stored in correspondence or a CRM.
- Introduce a second-approver rule — a payment above a set threshold, or to new account details, are approved by two different people.
- Set up liveness and anti-spoofing at the onboarding stage — don’t rely on visual review by a moderator.
- Add video stream injection detection — separately from basic liveness, since these are different technical layers of protection.
- Verify documents as a separate layer — a face matching the document photo does not replace checking the document itself for authenticity.
- Set up risk scoring and transaction monitoring — as a safety-net layer in case an attack gets past verification.
- Train employees to recognize signs of social engineering — urgency, secrecy, being told not to call back, a sudden change of communication channel.
- Don’t let an employee independently and instantly approve non-standard instructions — the right to pause and make a clarifying call must be formally established, not treated as a sign of distrust toward management.
- Limit the public exposure of top executives’ voices and video — the more material is publicly available, the easier it is to collect a sample for cloning.
Deepfakes have stopped being an experiment — in 5–7 minutes and for about 50 rubles, a fraudster can generate a video convincing enough to get through a video call wearing someone else’s face. Humans tell a forgery apart with 55.5% accuracy — that’s a coin toss, not protection. A business builds not one barrier but several: liveness and anti-spoofing at the verification stage, video stream injection detection, document verification, risk scoring, and mandatory payment confirmation via an independent channel. Liveness covers presentation attacks during video verification, but voice fraud, attacks via messaging apps, and synthetic video calls require the remaining layers — none of them solves the whole problem on its own.