Why Detectors Need to Train Against Their Own Weapons
Group-IB tracked over 8,000 deepfake injection attacks against loan-application liveness checks. The lesson isn't that detection failed once, it's that detection needs adversarial data to hold at all.
Deepfakes bypass liveness detection by feeding a photorealistic, AI-generated face into the video stream a system expects a live selfie from, using virtual camera software rather than holding a printed photo up to a real camera. The synthetic face blinks, turns, and responds to prompts on command, which is exactly what most liveness checks were built to confirm is present. The check passes because the system was never designed to ask whether the video feed itself is real.
Group-IB's fraud researchers counted 8,065 separate biometric injection attempts against loan-application liveness checks: fraudsters using virtual camera software to feed AI-generated deepfake video into the exact camera stream a bank's app expects a live selfie from. Each attempt is a small test of the same question: does this detector actually verify a live human, or does it just check that something moved on screen?
Most liveness systems were built to catch a printed photo held up to a webcam. That's a 2D problem. Today's attack is a photorealistic face that blinks on command, turns its head when asked, and responds to a liveness prompt in real time, injected through software the detector was never trained to look for. The fraud rate in these environments has grown accordingly: deepfakes now account for roughly 11 percent of global fraud activity, up from 7 percent two years earlier.
The instinct when a detector fails is to patch the specific hole: block the virtual camera driver, add a new liveness prompt, tighten the threshold. That buys weeks, not years. The attacker adapts to the patch faster than the patch gets shipped, because the attacker only needs one new technique while the defender has to anticipate all of them.
The alternative is to stop waiting for real fraud to teach the model what to look for. A detection model only gets as good as the hardest examples it's trained on, and real fraud attempts, by definition, are a sample of what already worked before someone caught it. That's a training set biased toward the attacker's past successes, not the attacker's next move.
Synthetic data flips that. Instead of harvesting fraud after the fact, a generator produces the adversarial material directly: identities that pass exactly the liveness and biometric checks a detector currently gets wrong, at whatever scale and demographic spread the real world hasn't yet supplied. Retrain the detector on that, and it's no longer learning from yesterday's incident report. It's learning from every plausible attack a generator can construct today.
This is also where the economics work in the defender's favor for once. The market for synthetic-data tooling itself is modest, but what it drives is not: data generation, model retraining, and serving all consume compute, and independent estimates put annual AI data-center spend at $1.2 to $1.7 trillion by 2030, with synthetic data projected to overtake real data as the primary training input across the industry. Hardening a detector isn't a side cost anymore. It's becoming the main workload.
TessLabs was built inside a live identity-verification business that spent eight years hardening detection against real fraud at population scale, then turned the generator that trained those detectors into a product of its own: mathematically locked synthetic identities that reproduce the exact edge case a model fails on, on demand.
See the Studio or read the white paper for the detection benchmarks.
Case study: Group-IB biometric injection-attack tracking, via Adaptive Security.
Measuring this on your own model
The first step is a sample built to your specification, which you score on your own detectors and benchmarks. No cost and no commitment.