The First Stage of KYC: Document Capture, Selfie, and Liveness as the Basis of Trust in the Session
A reliable KYC process begins with controlled data capture at the moment of the check. The user shows the document to the device camera, takes a selfie or a selfie with the document, then passes Liveness. This sequence gives the system the ability to assess the authenticity of the image source, the condition of the document itself, the correspondence of the face to the document owner, and the fact of a live person’s presence in the frame.
A scenario with the upload of pre-prepared files creates excess risk for the business, since the fraudster gains time to edit the image, substitute details, replace the photograph, generate a synthetic document, or prepare recaptured material. Photographing in the moment through the device camera reduces this risk, because the system receives data from a controlled stream and can verify the capture process itself, the frame quality, and the user’s behavior throughout the session.
Controlling Document Capture in the Moment
The first step of the KYC loop is confirming that the document is being shown to the device camera right during the session. At this level, the system controls the video-stream source and verifies that the image entering the loop comes from the physical camera of a smartphone or computer. For this, the integrity of the capture stream, the device characteristics, and signs of interference in the image-transmission channel are analyzed.
Such a check is necessary because there is a risk of bypassing even mandatory document photography. One possible scenario is the use of virtual cameras and the “injection” of frames into the video stream: instead of a real image from the camera, pre-prepared or generated visual content is fed into the KYC session. That is precisely why the protection includes a short but mandatory anti-spoofing stage at the capture-source level: the system must ensure that the frame was obtained from the device’s physical camera and that the video stream at the given moment has the signs of a live capture.
Verifying That the Original Document Is in the Frame
After confirming the stream source, the system determines the nature of the object in front of the camera. Its task is to establish that the user is showing an original document suitable for verification. At this stage, photographs of a phone, laptop, or TV screen, printouts, photocopies, recaptures from another medium, and synthetically generated images are identified.

Screen recapture is characterized by moiré, the display’s subpixel structure, a specific distribution of brightness and color, and features of the optical picture along the object’s edge. Printouts and copies are indicated by different surface textures, a different rendering of the security background, features of reflections, and the nature of the print layers.

Generative images give themselves away through artifacts in microtextures, security features, fonts, fine lines, stamps, and transitions between zones of the document. The combination of these features allows the system to establish that a physical original document, shown in a real environment, is present in the frame.
Frame Quality and the Document’s Suitability for Verification
The next layer of the KYC process is responsible for ensuring that the document can be reliably recognized and verified. The system determines the boundaries of the document in the frame and assesses the object’s area, position, perspective, sharpness, lighting, noise, and the presence of glare. If part of the document extends beyond the frame boundaries, the image is too dark or overexposed, important zones are covered by glare, and the text and security features lose their readability, the session receives a request for a repeat capture.
The result of this stage is an aligned and normalized image of the document, prepared for the subsequent checks. It is on this image that the system recognizes fields, isolates the owner’s photograph, analyzes the security background, reads the MRZ, and looks for stamps, signatures, and other mandatory elements.
Verifying Document Integrity and Traces of Tampering
Once the system has received a high-quality frame of an original document, analysis of the medium itself and its details begins.
Here the KYC loop looks for traces of physical and digital changes: data written in by pen, corrected digits, a re-glued photograph, stickers, erasures, painted-over zones, drawn-in elements, violations of the security ornament, texture anomalies, and local discrepancies in geometry.
Special attention is paid to the fields that fraudsters change most often: full name, date of birth, document number, expiry date, the owner’s photograph. The system verifies that the photograph is embedded in the document consistently with the background and security features, that the continuity of the pattern is preserved in the photo zone, and that the boundaries of the image and the surrounding areas look natural. For paper and plastic documents, stamps, signatures, seals, and other mandatory visual attributes are also analyzed. Their absence, anomalous position, a break in the template, or a mismatch with the document type raises the risk score.
OCR, MRZ, and the Logical Consistency of Fields
After the visual check, the system extracts the document’s details and conducts a logical validation. OCR reads data from the visible zone: full name, document number, date of birth, date of issue, expiry date, issuing authority, and other fields. Then these data are reconciled with the MRZ, if present, and with the template of the specific document type.
At this level, the format and length of the fields, the allowable characters, the check digits, the consistency of dates, the correspondence of the document number to the structure of the specific country, the presence of mandatory zones, and the mutual correspondence of the details are checked. Here too the system controls the completeness of the document’s composition: is there a photograph, is there a machine-readable zone, are the signature, seal, or stamp present if they are mandatory for this specimen. Such a layer makes it possible to identify forgeries that look visually convincing but contain internal contradictions.
A Selfie with the Document and a User Selfie as Confirmation of the Link Between the Document and the Person
The next stage links the presented document with the person undergoing verification. Depending on the scenario, the user may need to take a selfie with the document, a separate selfie, or pass a short face video-verification stage. The system obtains an image of the face against which the photograph in the document will be compared.
This stage is important for protection against the use of someone else’s documents and against scenarios where the genuine document of another person is presented. The KYC platform isolates the face on the document, extracts biometric features from the owner’s photograph, and matches them against the user’s face in the selfie or the video stream. The result is an assessment of the biometric correspondence between the document and the person undergoing onboarding.
Liveness as Confirmation of a Live Person’s Presence
After the selfie stage, the system must confirm the presence of a live person in front of the camera. For this, Liveness is applied. The check assesses the signs of a real face in the moment: natural movements, the depth and shape of the face surface, reactions to the capture scenario, skin texture, frame dynamics, and other parameters available to the specific implementation.
Liveness closes off a separate class of attacks: showing a photograph, a screen, a video recording, a deepfake, or a presentation attack. In combination with control of the video-stream source, this stage forms a stronger defense of the KYC session. One layer confirms the trustworthiness of the image-capture channel, the other layer confirms the presence of a live person in front of the camera.
The Final Logic of the Stage
Thus, the first major block of the KYC process is built as a sequence of interconnected checks. The system must confirm that the document is being captured by the device camera in the moment, that the original document is in the frame, that the image is of sufficient quality for analysis, that the details, stamps, signatures, and security features look consistent, that the user’s face matches the photograph in the document, and that a live person is present in front of the camera.
Only with such a sequence does the subsequent KYC decision gain a solid foundation. Each layer reinforces the next: capture-source control increases trust in the frame, document verification increases trust in the details, face match links the document to the person, and Liveness confirms the reality of the user’s presence in the session.
Traces of Editing in the Document Image
Passing the input control of format, metadata, and capture quality does not guarantee that the content of the document has not been changed. The next level is the search for traces of pixel-level editing. The task is to detect areas that were inserted, redrawn, cloned, or otherwise modified after the original image was obtained.
A forger working in a graphics editor strives to make the introduced changes as imperceptible to the eye as possible. However, every action — whether inserting part of another image or manually correcting an individual digit — leaves behind specific statistical and geometric traces (artifacts). These artifacts, which are often indistinguishable to a human, are reliably recorded by digital image forensics algorithms. Below are five main categories of such artifacts.
Local Insertions and Splices by Inconsistency of Textures and Noise
The most common type of forgery is inserting a fragment from another image: someone else’s photo, surname, date of birth, number. Technically this is a splicing operation: an area from one source is placed into the target image. Even with careful matching of scale and color, the insertion produces diagnostic signals.
| Category | Description |
|---|---|
| Inconsistency of the noise model | Any camera or scanner introduces characteristic sensor noise whose intensity and structure depend on the device model, ISO, exposure, and processing pipeline. A fragment from another device or after different post-processing differs in the local noise statistics. Algorithms for estimating the noise level function (NLF) compare the dependence of noise variance on brightness across different blocks of the image. An area whose NLF curve statistically deviates from the global one is flagged as suspicious. For documents the method is especially effective: uniform background zones (fields, substrate) give a stable noise estimate, and a local deviation is well distinguishable. |
| Texture breaks at the insertion boundaries | When two fragments are combined, a transition zone inevitably arises in which the paper microtexture, print graininess, or the security-grid pattern loses continuity. Convolutional neural networks trained on «original — forgery» pairs record such boundary breaks through the analysis of gradients and high-frequency components. Modern architectures use multi-scale attention maps that take into account both fine texture anomalies at the pixel level and larger semantic inconsistencies — a break in a security line or a shift in line markup. |
The insertion detector works pixel by pixel or block by block: for each area, the probability of belonging to an edited region is computed. The result is an anomaly map (heatmap), which is passed to the overall decision-making loop.
Cloning and Erasing of Areas in the Background and Fields
If splicing involves introducing material from another source, then cloning (copy-move) is the transfer of a fragment within the same image. A typical scenario: the forger copies a section of clean background and overlays it onto a stamp, record, or mark. Another variant is erasing (inpainting): removing an object with automatic filling of the «reconstructed» content.
Clone detection. Algorithms identify areas in the image with anomalously high similarity not explained by the document’s natural structure. The classic approach: the image is divided into overlapping blocks, invariant descriptors are computed for each, then pairs of blocks with minimal distance are sought. Neural-network methods solve the task «end-to-end»: one branch detects anomalies by visual artifacts, the second by the similarity of regions, and a fusion module forms the final forgery map. Such detectors are resilient to typical masking attempts: a slight rotation, scaling, a change in brightness, and light blurring of the cloned area.
Erasing detection. Erasing is harder to detect: modern inpainting algorithms generate visually plausible filling. Traces remain: anomalous blurring in the fill zone, violation of regular patterns (ruling lines, the security grid, guilloche patterns), and a mismatch of texture statistics with the surrounding areas. For documents this is especially relevant: security features (microprint, ornament, colored fibers) create predictable regular structures, and their local violation serves as a reliable indicator of tampering.
Non-Uniform Compression and Re-Saving Artifacts
The overwhelming majority of document images are stored in JPEG format. Each save introduces characteristic artifacts: the image is divided into 8×8 pixel blocks, and the discrete cosine transform (DCT) coefficients are quantized by a given matrix. This process leaves a measurable «fingerprint» — a specific distribution of DCT coefficients and block boundaries. If the entire image was compressed uniformly, the fingerprint is uniform. If part is edited and inserted, a mismatch in compression history arises.
Double JPEG compression. When a JPEG is opened in an editor, changes are made, and it is saved, the edited areas undergo one compression cycle while the untouched ones undergo two. This creates statistically distinguishable distributions of DCT coefficients: in doubly compressed blocks, periodic artifacts appear (the double quantization effect, the DQ effect). Detectors analyze the histograms of DCT coefficients block by block and classify each block as singly or doubly compressed. An area where the nature of compression differs from the rest of the image has a high probability of having been edited.
Error Level Analysis (ELA). The image is re-saved with a deliberately known JPEG quality coefficient, after which the pixel-by-pixel difference between the original and the recompressed copy is computed. Uniform areas give an even error map, while areas with a different compression history stand out with an anomalous level of residual error. ELA is a fast heuristic filter: it does not give an unambiguous answer, but it reliably indicates zones for detailed analysis. For documents the method works well on contrasting elements — text and field boundaries, where JPEG artifacts are most pronounced.
Redrawing of Text and Digits in the Details
Changing specific characters — the passport number, date of birth, name — is one of the most frequent targets of forgery. A forger may paint over the original characters and draw new ones, use the «stamp» tool to transfer digits from another part of the document, or overlay a text layer on top of the original.
Font anomalies. State-issued documents are printed on standardized equipment with a fixed set of typefaces and print parameters. A redrawn character almost always differs in kerning, stroke thickness, baseline, or degree of rasterization. OCR verification matches the metrics of the recognized characters against the reference font parameters for the given document type. An anomalous character gets a low confidence score — a signal for an additional check.
Texture and compression traces of substitution. A redrawn area differs in microtexture: different graininess, the absence of the edge irregularity of letters characteristic of printing, a different distribution of compression artifacts. Neural-network detectors trained on datasets of document forgeries (DocTamper, CVPR 2023, 170,000 images) analyze the visual consistency of characters in context: the network identifies areas where the rendering of a character is statistically incompatible with its surroundings.
Boundary artifacts. During insertion or redrawing, a thin «halo» often remains around the character — an area where the background pixels were smoothed or interpolated. It may be invisible on the screen but is recorded during the analysis of high-frequency components or during an ELA check.
Inconsistency of Color, Lighting, and Shadows
Transferring a fragment from one image to another rarely preserves full color and light unity. Differences in shooting conditions — the angle of lighting, color temperature, white balance — lead to mismatches imperceptible during a cursory review but algorithmically detectable.
Color uniformity. The substrate of a genuine document has a unified color profile determined by one light source and one exposure. An inserted fragment introduces a local shift in the color channels. Algorithms analyze the color distribution in spaces resilient to brightness variations (Lab, HSV) and identify statistically incompatible areas.
The direction and nature of lighting. On a document with side or uneven lighting, soft brightness gradients and micro-shadows form near relief elements (embossing, raised print, laminate). A fragment with a different light direction has differently oriented shadows and highlights. Methods for estimating the lighting direction build a model of the light field for the entire image and detect areas where the local lighting vector conflicts with the global one.
Chromatic aberration. A camera lens introduces a color shift that depends on the distance from the optical center of the frame. A moved or borrowed fragment does not match the aberration model calculated for the given lens and position in the frame. This signal complements the noise and texture analysis, forming another independent detection channel.
The combination of checks — noise analysis, clone detection, compression analysis, font verification, and the check of color-light consistency — forms a multi-signal anti-fraud layer that works before OCR and structural checks are engaged. Each method individually can produce false positives: a shadow on a document fold can imitate a color anomaly, and low scan quality can smooth out double-compression artifacts. A reliable result is ensured by aggregating many weak signals into a single risk score — more on this in the section devoted to the final scoring.
Authenticity Verification by OCR and MRZ with Consistency Rules
Visual analysis of the image reveals traces of retouching and montage, but pixel artifacts alone are insufficient for a confident verdict. The next line is verifying the semantic content of the document: do the extracted data conform to the format, the internal rules, and each other. Here OCR, the machine-readable zone (MRZ), barcodes, and reference templates come into play. Each check operates on structured information that the system reads from the document and reconciles against a set of formal and logical rules.
The principle: a genuine document is consistent across all data layers. The surname in the visual zone matches the surname in the MRZ, the date of birth in the barcode corresponds to the date in the text fields, the security background matches the reference for the given type and series. Any inconsistency is a signal requiring assessment.
Verifying Document Integrity by OCR Results
OCR extracts textual data from the document’s visual zone (VIZ — Visual Inspection Zone): full name, date of birth, document number, address, date of issue, and expiry date. The anti-fraud system’s task begins where the recognized text passes through a chain of validations.
| Category | Description |
|---|---|
| Field format | Each document type has rigid rules: the length of the passport series and number, the allowable date ranges, the set of characters in specific positions. If OCR read the number of a Russian passport as «12 34 5678A9», the presence of a letter in the numeric series is a formal violation of the structure. The system records the discrepancy automatically. |
| Cross-validation of fields within the document | The date of birth precedes the date of issue; the date of issue is earlier than the expiry date; the age at the time of issue falls within the acceptable range for the given document type. The department code in a Russian passport is linked to the region, and this link is verifiable against open reference books. A mismatch between the department code and the region of issue is one of the common signs of a homemade forgery. |
| Font and positional analysis | Advanced OCR systems record the typeface, point size, and inter-character spacing. Genuine documents are printed on industrial equipment with fixed typographic parameters. Detecting a font in one field different from the others, or a deviation of the line spacing from the norm for the given template, is a signal for an in-depth check. |
Reconciling the Layout and Security Background Against the Reference Document Template
Each type of identity document has an approved design: the placement of fields, the size of the photograph, the color scheme, the ornamental substrate. The anti-fraud system stores reference templates — descriptions of the layouts and security features for each supported type and series.
The geometric check reconciles the coordinates of key zones: the position of the line with the document number, the area for the photograph, the position of the MRZ. A shift of the fields by a few millimeters relative to the reference may indicate a manual reassembly of the layout in a graphics editor.
The security background — the guilloche grid (guilloché) — is one of the most reliable indicators of authenticity in a digital check. The guilloche pattern is a system of thin curved lines built according to mathematical rules. Its complexity makes exact reproduction extremely labor-intensive even with professional graphics software. Neural-network models compare fragments of the presented document’s background with the reference pattern and record distortions: line breaks, violations of periodicity, deviations of the color gradient. Research by Al-Ghadi et al. (2022) on the MIDV-2020 and FMIDV datasets confirms that CNN models based on contrastive and adversarial learning distinguish a genuine from a forged guilloche pattern with F1-scores from 75 to 100% depending on the document type and configuration parameters.
Reconciliation against a template is effective to the extent that the base of references is complete and up to date. Documents are updated: series change, new security features are introduced, the design is adjusted. Regular replenishment of the template base and tracking of changes in the regulations of issuing countries is a mandatory condition for the operability of this verification layer.
Verifying the MRZ by Format and Check Digits
The machine-readable zone (MRZ) is a standardized data block printed in OCR-B font at the bottom of the page with personal data. The MRZ format is defined by the ICAO Doc 9303 standard and describes three main formats: TD1 (three lines of 30 characters, ID cards), TD2 (two lines of 36 characters), and TD3 (two lines of 44 characters, passports).
The MRZ allows only uppercase Latin letters A–Z, digits 0–9, and the filler character «<». Any other character in the read line is an unambiguous indicator of a recognition error or an integrity violation.
The key protection mechanism is check digits. Each character is converted into a numeric value (digits retain their value, letters A–Z are encoded by the numbers 10–35, «<» equals 0), the sequence is multiplied by the repeating weighting coefficients 7, 3, 1, the results are summed, and the remainder of dividing the sum by 10 gives the check digit.
In the passport MRZ (TD3), the check digits protect: the document number, date of birth, expiry date, personal number (if present), and also a composite check digit covering several fields of the second line. If even one character in a protected field is changed, the recomputed check digit will not match the specified one.
Verifying the MRZ takes milliseconds, requires no access to external databases, and is performed entirely locally. A successful check of the check digits confirms the internal consistency of the MRZ but does not prove the authenticity of the document: a fraudster who possesses the algorithm can generate an MRZ with correct check digits for fictitious data. MRZ verification is a necessary but not sufficient element of verification, and it is always applied in combination with other methods.
The ninth edition of ICAO Doc 9303, which took effect on January 1, 2026, introduces standardized two-letter document-type codes in the MRZ: «PP» for an ordinary passport, «PD» for a diplomatic one, «PE» for an emergency one. From this date, passports using the second letter in the type code are required to apply the standard designations from Doc 9303-4; from January 1, 2028, the two-letter code will become mandatory for all newly issued passports. Anti-fraud systems must account for the transition period: documents of new series are checked by the updated rules, documents of previous ones by the former rules.
Cross-validation of fields between the MRZ, VIZ, and barcodes reveals forgeries that cannot be detected by pixel analysis alone — but only provided that the reference base is up to date and the transliteration rules are taken into account. The NeuroVision AI-OCR module extracts data from the visual zone with 99.85% accuracy for printed documents, simultaneously parses the MRZ per the ICAO Doc 9303 standards with a recalculation of check digits, and reconciles the fields character by character, taking into account the transliteration tables of issuing countries. The platform’s coverage is more than 10,000 document types, and the references are updated when new series are issued and regulations change. We will audit your current document flow, determine which types and regions require priority coverage, and propose a configuration with an optimal balance of automation and manual control.
Searching for Discrepancies Between the MRZ and Visual Details
Cross-reconciliation of the MRZ and the visual zone (VIZ) is one of the most effective methods for detecting forgeries at the level of text or image editing. An attacker who changed the full name or date of birth in the visual part often leaves the MRZ untouched — out of carelessness or because they do not possess the check-digit recalculation algorithm.
The system extracts the same fields from two sources — the VIZ (via OCR) and the MRZ (via a specialized parser) — and reconciles them character by character, taking into account transliteration rules. The ICAO Doc 9303 standard, part 3, defines the order of transliteration: the German «ü» is rendered as «UE», the Icelandic «ð» as «D». For Cyrillic names, the transliteration tables of the issuing country are applied. A mismatch not explained by transliteration rules is a weighty alarm signal.
Typical discrepancies during cross-checking: the surname or first name in the VIZ does not match the MRZ (the most frequent case when substituting text fields in an editor); the date of birth differs; the document number in the MRZ does not match the one printed in the visual zone; the citizenship code in the MRZ does not correspond to the country indicated in the VIZ; the sex in the MRZ (M/F/«<») contradicts the data of the visual zone.
Each discrepancy receives a weighting coefficient. A document-number mismatch is scored higher than a discrepancy in a single character of the surname, which may be explained by the variability of transliteration. The set of discrepancies is passed to the risk-aggregation module for routing: automatic rejection, manual review, or a request for additional documents.
Verifying Barcodes and the QR Code for Consistency with the Fields
A number of identity documents contain machine-readable elements besides the MRZ: two-dimensional PDF417 barcodes (driver’s licenses, primarily in the US and Canada), QR codes (biometric ID cards of a number of EU countries), linear barcodes on some types of visas. These elements encode the owner’s personal data — often in a volume comparable to the MRZ or exceeding it.
The system decodes the content of the barcode or QR code and reconciles the extracted fields against the data from the VIZ and MRZ. A match of all three sources is a strong positive signal. A discrepancy in even one field is grounds for raising the risk level.
PDF417 uses Reed–Solomon error correction: data damaged during printing or capture is restored within certain limits. If the code structure is violated beyond the corrective capacity or does not conform to the expected specification for the given document type, this is also an indicator of a problem.
The presence of a correct barcode does not by itself guarantee authenticity. The PDF417 format is open and documented (ISO/IEC 15438), and an attacker is capable of generating a barcode with arbitrary data. A number of countries solve this problem with a cryptographic signature inside the code: the data is signed with the issuer’s private key, and the verifying party validates the signature with the public key.

This approach is implemented, in particular, in French biometric ID cards, where the QR code contains an electronic seal (cachet électronique visible — CEV) per the ISO 22376 and ISO 22385 standards, which makes it possible to confirm the authenticity of the data without access to a centralized database.
For documents without a cryptographic signature, the barcode is an additional consistency layer: it cannot be the sole proof of authenticity, but a discrepancy with the VIZ or MRZ is a reliable indicator of manipulation.
The checks of OCR, MRZ, templates, and machine-readable codes form the logical validation layer. It works not with pixels but with meaning — and it reveals forgeries carried out with technical precision but containing internal contradictions that cannot be eliminated without a full understanding of the structure and rules of the specific document.
Detecting Substitution of the Owner’s Photograph

The photograph is the most vulnerable element of a document from the standpoint of forgery. By replacing the image, an attacker turns someone else’s genuine document into their own: the details, serial numbers, MRZ, and security features remain genuine, while the substituted photo makes it possible to pass verification under someone else’s name. According to ICAO and Europol, photo substitution is among the most common ways of forging passports and identity documents. Analysis of the photo zone must work as a standalone verification loop rather than merely a supplement to OCR and MRZ validation.
Analyzing the Photo Zone for Signs of Paste-In or Replacement
In the case of physical replacement of the photo in a paper document (with subsequent scanning), traces remain that are distinguishable in the digital image. In the case of digital forgery — substitution in a graphics editor — the nature of the artifacts is different, but the detection principle is common: the system looks for local inconsistencies that do not occur in the original.
The first thing the algorithm pays attention to is the boundary of the photograph. In a genuine document, the transition from the image to the page background has a characteristic texture: for paper passports this is an even line with the same level of compression artifacts on both sides, for polycarbonate cards it is a uniform structure without visual breaks. A physical paste-in leaves a thin shadow strip along the edge on the scan, a micro-shift of the plane (noticeable by a difference in sharpness), or a thickening of the substrate. Digital replacement gives itself away by a difference in JPEG compression levels (ELA), a mismatch of the noise profile, or a sharp jump in color temperature.
The system also analyzes the internal uniformity of the photo zone, comparing the noise statistics and the brightness and color histograms with the rest of the document page area. In the original, the parameters are consistent, since the document was made on a single piece of equipment. An inserted image almost always has different statistics: a different sensor, a different exposure, a different post-processing algorithm.
According to IEEE publications, the combination of ELA with a CNN achieves an accuracy of 94–96% on public benchmarks, although real-world effectiveness depends on the quality of the input image and the diversity of the training set.
A separate class of attacks is face morphing. An attacker generates a hybrid image combining the features of two people. Such an image can pass an algorithmic check for both faces simultaneously. The NIST publication — NISTIR 8584 (August 2025) — examines methods of morph detection and notes: the best algorithms detect up to 100% of morphs at an FAR of 1%, but only if they are trained on examples from the same generator. On unfamiliar generators, accuracy can drop below 40%. Models need regular updating of their training data as new synthesis tools appear.
Consistency of the Ornament and Security Features Around the Photo
Modern documents are designed so that replacing the photograph inevitably violates the integrity of the security features. On the data page of a passport, the ornamental pattern, the guilloche grid, or the microtext substrate runs continuously through the photo zone and beyond it. If the image is replaced, the lines of the pattern break, shift, or disappear at the boundary of the insertion.
For polycarbonate documents (ID cards, new-generation biometric passports), the protection is even more serious. The photo is personalized by laser engraving inside the polycarbonate layers, with a transparent holographic film with an optically variable device (OVD) applied on top. Delamination of such a construction destroys both the engraving and the hologram. In the digital image, this manifests as the absence or distortion of the holographic reflection in the photo zone, a violation of the continuity of the security pattern, a mismatch in the position of the OVD elements relative to the face.
Automated verification is built on comparison with the reference template of the specific type and series of the document. The system knows what the security background looks like, where the ornament lines run, and what the geometry of the holographic elements is. Deviations are recorded through a pixel-by-pixel comparison of the background structure in the photo zone and beyond it: a break in the ornament at the boundary of the image, a change in the pitch, the angle of inclination, or the color hue is a signal of tampering. Accuracy depends on the completeness of the reference base: industrial verification platforms work with bases covering thousands of document types from dozens of countries and regularly update the references when new series are issued.
Holographic and OVD elements appear differently depending on the shooting conditions. In a document photograph under certain lighting, the hologram appears as a bright glare, and its presence in the correct position serves as additional confirmation of authenticity. The absence of the expected glare or its uncharacteristic shape is another indicator for the anti-fraud engine.
Matching the Face on the Document Against the Selfie in KYC
Even if the photograph in the document is not replaced, this does not guarantee that the document is being presented by its lawful owner. The final line of KYC is matching the face from the document photo against a live image (selfie) or video frame of the user.
Technically this is a one-to-one verification task: the algorithm extracts a feature vector (embedding) from the document photograph and from the selfie, then computes a similarity metric. If the value is above the threshold, the identity is confirmed. Leading algorithms tested by NIST in the FRTE (Face Recognition Technology Evaluation) program demonstrate a false non-match (FNMR) at the level of fractions of a percent at a fixed false-match threshold (FMR) of 0.01%. In real conditions, accuracy noticeably depends on data quality: poor lighting, low selfie resolution, age-related changes, a beard, glasses, makeup — all of this increases the share of false rejections. The similarity threshold is a compromise between security and conversion, and its value is configured for the specific business scenario.
Face matching by itself does not protect against presentation attacks: an attacker can hold a printed photo up to the camera, play back a video recording, or use real-time deepfake generation. Face matching works in combination with liveness detection (a «vitality» check), which confirms that a real person is in front of the camera. Modern liveness modules analyze micro-movements, skin texture, the reaction to light stimuli, the optical properties of the 3D surface of the face, and other features that cannot be reproduced by a flat medium. Certification against ISO/IEC 30107-3 (levels 1 and 2, iBeta testing) confirms the implementation’s resilience to the main types of attacks — photo, video, masks.
A separate risk is injection attacks, when an attacker substitutes the video stream at the software level, bypassing the device camera. Protection is built on controlling the integrity of the SDK and the data-transmission channel: the system verifies that the image was obtained specifically from a physical camera rather than from a virtual source.
The results of face matching, liveness, and analysis of the photo zone are aggregated into a single signal. A low similarity score, a failed liveness check, or the detection of artifacts in the photo zone — each factor raises the final risk score. At a threshold value, the session is routed to manual review or automatic rejection. The multi-level structure reduces the number of false positives without harming protection: each loop covers the weak points of its neighbor.
The multi-level verification structure — document, photo, liveness, cross-validation — works only when all loops are integrated and exchange signals. The NeuroVision platform combines in a single pipeline AI-OCR with anti-fraud checks, biometric face verification (Enface accuracy — 99.74%, TOP-1 among Russian algorithms in NIST testing), a passive liveness check with 99.9% accuracy, and aggregated scoring with explainable reasons for the decision. The cost of the full cycle is from 35 to 50 rubles per check depending on the set of modules and volume, and the face processing time is under 0.1 second. Deployment is possible in the cloud, within your perimeter, or in a hybrid format; the launch benchmark is 3 to 7 days including scenario configuration and information-security requirements. We will calculate the cost for your application volume and select a module configuration that will satisfy the regulatory requirements of your jurisdiction.
The Anti-Fraud Decision and Detection Quality Control
Each individual check — texture analysis, MRZ reconciliation, clone detection, lighting inconsistency — gives a local signal. By itself it is rarely sufficient for a categorical conclusion: a compression artifact may turn out to be a consequence of conversion, and a font discrepancy the result of a legitimate form update. The anti-fraud system’s task is to assemble the disparate signals into a single decision, explain its logic to an operator or auditor, set clear boundaries of automation, and ensure the stability of quality over time.
Aggregating Signals into a Risk Score and Explainable Reasons
The risk score is a numerical estimate of the aggregate probability that the document is forged, edited, or does not belong to the applicant. Dozens of parameters come to the input: the results of the analysis of noise and compression artifacts, the degree of match between OCR and MRZ, the state of the security features, the metric of comparing the face with the selfie, the file metadata, signs of capture from a screen.
Each parameter undergoes normalization and receives a weight. The weighting coefficients can be set by expert rules, a statistical model, or an ensemble of both approaches. The most robust results are given by a hybrid architecture: rules fix hard constraints (an invalid MRZ check digit — immediate rejection), while the ML model evaluates the combination of soft signals, each of which is not critical on its own but in combination indicates a forgery.
The final score is not just a number. The regulatory requirements of a number of jurisdictions (in particular, Article 22 of the GDPR and similar norms on automated decision-making) require providing the data subject with an explanation of a decision that significantly affects their rights. Explainability is implemented through a list of reason codes: the specific detector that fired, its contribution to the final score, the threshold at which the signal is considered significant. For each check, a structured response is formed: the score, the status, the list of reasons with priorities and, if necessary, visual markers on the image — the zones that caused the trigger.
This approach solves two tasks. The manual-review operator receives a concrete analysis route: which zones of the document to check first and which discrepancies were found. The compliance service can demonstrate to an auditor or regulator that the decision was made on the basis of verifiable criteria rather than an arbitrary «black box» assessment.
Decision Thresholds and Routing to Manual Review
The aggregated score divides documents into three categories: automatic approval, automatic rejection, and a zone of uncertainty requiring manual verification. The boundaries are set by a pair of thresholds — a lower one (below which the document is considered trustworthy) and an upper one (above which a forgery is determined with high confidence).
The choice of thresholds is a manageable compromise between the share of false positives (FPR), the share of missed forgeries (FNR), and the volume of manual reviews. Tightening the upper threshold reduces FNR but increases the share of cases in manual review and, consequently, operational costs. Softening the lower one speeds up onboarding but raises the risk of missing a forgery.
Thresholds are configured for each document type, region of issue, and business scenario separately. A passport with a well-standardized form and MRZ allows stricter automatic rules than an income statement without security features. In high-risk scenarios (opening a bank account, issuing a loan), the acceptable share of missed forgeries must be minimal, even at the cost of increasing the manual flow. In low-risk ones (age verification, address confirmation), more automatic decisions are acceptable.
Routing is supplemented by prioritization. Cases with the largest number of triggered detectors or anomalous combinations of signals enter the queue first. This allows operators to focus on genuinely suspicious applications rather than spending resources on borderline cases, which more often turn out to be legitimate.
A separate element is feedback from manual review. Operators’ decisions (confirmed / rejected) return to the system as labeled data and are used to calibrate thresholds: if operators consistently confirm cases of a certain type, the threshold for this category can be adjusted, reducing the manual flow without an increase in risk. A closed feedback loop is one of the key mechanisms for improving the effectiveness of the anti-fraud decision over time.
Document Verification Quality Metrics and Target Error Levels
To assess the operation of an anti-fraud system, overall accuracy is insufficient: with a high share of legitimate documents in the flow, even a primitive model that approves everything indiscriminately will show a formally high percentage but will miss every forgery. The key metrics are sensitive to errors of each type.
| Category | Description |
|---|---|
| False Positive Rate (FPR) | The share of genuine documents erroneously rejected. Each false rejection is a lost customer, a support inquiry, a blow to conversion. The target FPR depends on the context: in mass online onboarding — 1–3%, in premium services with a high acquisition cost — below 1%. |
| False Negative Rate (FNR) | The share of forged documents erroneously accepted. A direct financial and compliance risk. In the banking sector and cryptocurrency services, the benchmark is values below 1%, and in some cases below 0.1%. |
| Automation Rate | The share of applications processed fully automatically. Typical benchmarks for mature solutions are 85–95%, but the specific value depends on the quality of the input flow, the diversity of document types, and the threshold settings. |
In addition to the main metrics, the following are tracked: processing speed (latency), time to the final decision including manual review, precision and recall for individual types of forgery. A system may demonstrate acceptable aggregated indicators but miss a specific class of attack — the redrawing of digits or the substitution of a photograph in a certain way.
Metrics are recorded in the breakdown of document types, countries of issue, intake channels, and time periods. Without such detail, degradation on a narrow segment remains unnoticed against the background of the overall averages. Regular reporting is a mandatory element both of internal control and of interaction with the regulator, which has the right to request data on the accuracy of automated decisions.
Monitoring Drift and Updating Rules and Models
An anti-fraud system operates under conditions of constant change. Fraudsters adapt their methods: instead of crude paste-in they move to generative models, instead of editing a JPEG they move to recapture from a screen, instead of forged passports they move to less protected document types. In parallel, legitimate data changes: forms are updated, new series appear, the geographic and demographic composition of the flow shifts. A model trained on yesterday’s data and not adapted to the current reality inevitably loses quality.
This phenomenon is drift: a change in the distribution of the input data (data drift) or in the relationship between features and the target variable (concept drift). Data drift manifests when a new document specimen that was not in the training set appears en masse. Concept drift is when a previously reliable sign of forgery stops working because fraudsters have learned to bypass it.
The monitoring system tracks several levels. At the level of the input data, the distributions of key features are monitored: a sharp change in the share of documents of a certain type or region is a signal for analysis. Statistical tests (the Kolmogorov–Smirnov criterion for continuous features, the Population Stability Index — PSI) formalize the significance threshold of a deviation.
At the model level, the distribution of the output scores and the share of cases in the uncertainty zone are tracked. A rise in this share indicates that the model is «losing confidence» — the input data increasingly falls outside its training experience. At the business-metrics level, FPR, FNR, the share of manual reviews, onboarding conversion, and the number of confirmed fraud incidents are monitored. A deterioration in any indicator is a trigger for investigation.
The response to drift depends on its nature. The appearance of a new form requires updating the reference templates and validation rules, but not retraining the ML models. A change in fraudster tactics may require retraining on fresh labeled data with samples of new types of forgery. In critical cases, an interim measure is tightening the thresholds: more cases go to manual review, which reduces FNR until an updated model is released.
Data drift and a change in falsification tactics devalue even a precisely configured system — without continuous monitoring of metrics and updating of references, degradation is detected only by its consequences. NeuroVision takes on the maintenance of the anti-fraud loop: monitoring FPR, FNR, and the share of automatic decisions in the breakdown of document types and regions, updating reference templates when new series are issued, retraining models on fresh data with validation in shadow mode. The platform availability SLA is 99.99%, the trial period is up to 1 month, and you are assigned a personal account manager and 24/7 technical support. We will start with a joint elaboration of your scenarios and target quality metrics — as an output, you will receive an agreed-upon launch plan with clear acceptance criteria at each stage.
The update cycle includes collecting new labeled data (including from operator feedback), retraining the model, validation on a test set, a parallel launch (shadow mode) alongside the production version, and switching upon confirmation of improved metrics. Versioning of models and rules with the ability to roll back is a mandatory condition: if a new version degrades quality on some segment, the system must allow a quick return to the previous one.
The frequency of updates is determined by the speed of changes and the criticality of the task. In high-load KYC services processing hundreds of thousands of applications a month, rules and templates are adjusted as new specimens arrive, and models undergo revision monthly or quarterly. For less dynamic scenarios, a quarterly or semi-annual cycle is sufficient, but the monitoring of metrics must remain continuous. Without it, an anti-fraud solution degrades unnoticed, and the consequences are discovered only in the form of missed incidents or a surge of false rejections.
Detecting forgeries requires an established chain: from byte-level validation of the file and analysis of pixel artifacts to cross-checking of MRZ, OCR, barcodes, and biometrics. No single layer is self-sufficient — a crude paste-in will be missed by noise analysis without reconciliation against a reference template, and correctly recomputed MRZ check digits will not save against photo morphing. Aggregating many weak signals into a single risk score with explainable reasons makes it possible to maintain the balance between missed forgeries, false rejections, and the volume of manual reviews.
The KYC process remains a living system: forms are updated, falsification tactics grow more complex, the distribution of the input flow shifts. Drift monitoring, a closed feedback loop from operators, and regular updating of models and templates are conditions without which even a precisely configured solution loses quality over time. Understanding the verification architecture at each level gives the implementation team clear evaluation criteria: which signals should be extracted, how they are combined, and which metrics confirm that the system is working.