The Synthetic Biometric Data Landscape: Who's Building What
Synthetic biometric data is a named, growing category with a small number of real players. Here's who's building what, and where each one actually sits in the stack.
Synthetic biometric data is a real, named category, not a research curiosity, but the number of companies actually building it is small and the players don't all do the same job. Some generate synthetic faces for computer vision broadly. Some generate synthetic tabular data for regulated industries and happen to touch identity as one use case among many. Only a few build biometric-specific synthetic data paired with the detection capability to use it defensively.
Who's in the category
Precise Biometrics, a Swedish biometrics and cybersecurity firm with a long history in fingerprint and palm data, added synthetic biometric data generation as a formal product line in 2025, addressing what it calls a scarcity of quality synthetic biometric data for its own anti-spoofing and liveness work and for customers.
Datagen, an Israeli company acquired into Nvidia's ecosystem, generates synthetic faces and body poses for computer vision broadly, human-centric datasets used across face recognition, driver monitoring, and related applications, built for volume and diversity rather than for detection-hardening specifically.
Synthesis AI builds synthetic humans using GANs and CGI, with FaceAPI as an early product aimed at facial verification, teleconferencing, and driver monitoring. Like Datagen, it sits closer to general computer-vision data supply than to a dedicated fraud-detection use case.
Mostly AI, based in Austria, generates synthetic tabular data for regulated industries like banking and insurance. It touches identity data as a category, not biometric imagery specifically, and its core differentiator is privacy-preserving structured data rather than faces or voices.
TessLabs differs from all four on two counts. First, it's deterministic rather than purely generative: every identity is a fixed mathematical seed, reproduced exactly, with individual attributes swept independently rather than resampled from a distribution, which is what makes a specific edge case reproducible on demand. Second, it wasn't built as a synthetic-data product first. It grew inside a live, enterprise-grade identity-verification business handling roughly 15 million real onboardings over eight years, where the generator existed to train the discriminator against real fraud at population scale before either became a standalone offering.
Why the distinction matters to a buyer
A generator built for computer-vision diversity and a generator built to train a fraud detector against adversarial edge cases solve different problems, even though both produce synthetic faces. The first optimizes for breadth and realism. The second optimizes for reproducibility, calibration to specific fraud-exposed populations, and pairing with a discriminator that gets measured against real detection benchmarks rather than just visual fidelity.
That difference shows up directly in due diligence. A synthetic-data company being evaluated as an acquisition target or a vendor should be able to answer three questions plainly: can it reproduce the exact same identity on demand for a controlled attribute sweep, was the generator ever tested against real fraud rather than just visual quality metrics, and does it ship a detector alongside the generator or only the generator on its own.
FAQ
Is there more than one company building synthetic biometric data?
Yes, though the category is small. Precise Biometrics, Datagen, Synthesis AI, and Mostly AI all touch synthetic identity or biometric data in some form, alongside dedicated players like TessLabs, but they differ significantly in whether they're deterministic or purely generative, and whether they pair the generator with a trained detector.
What's the difference between a generative and a deterministic synthetic-face system?
A generative system samples a new face from a learned distribution each time, so the same identity can't be reproduced on demand. A deterministic system fixes each identity to a mathematical seed, so the identical identity can be reproduced exactly while individual attributes are varied independently, which matters for isolating specific edge cases a detector fails on.
Why would an acquirer care about a synthetic-data company's detection track record, not just its generator?
A generator that's only ever been measured on visual realism hasn't been proven against the thing that actually matters commercially, which is whether it can train a detector to catch real fraud. A generator-and-detector pair with a documented benchmark against production fraud is a materially different asset than a generator alone.
Read the white paper to see TessLabs' detection benchmarks, or book a call to discuss the category in more depth.
Sources: Precise Biometrics synthetic biometric data launch via Biometric Update and company release; Datagen and Synthesis AI overview via Turing Post; Mostly AI via CB Insights.
Measuring this on your own model
The first step is a sample built to your specification, which you score on your own detectors and benchmarks. No cost and no commitment.