What Is Synthetic Media?
Synthetic media is image, video, audio, or text content that algorithms, most often artificial intelligence, have generated or substantially altered. Deepfake media is a subset of synthetic media: AI-generated or manipulated content that depicts a real, identifiable person or event convincingly enough to pass as authentic.
The term describes how algorithms made or altered the content, and the same models that produce a cloned executive voice also produce a voice bank (opens in new tab) for a person losing their speech. Consent, disclosure, and intent separate legitimate use from harmful use.
How synthetic media works
Generative models learn the statistical patterns of real media, then produce new output that matches those patterns. Generative adversarial networks pit a generator against a discriminator that judges (opens in new tab) it for authenticity, and they drove most video face-swap tooling through about 2023.
Classic face-swap deepfakes train two autoencoders with a shared encoder to reconstruct a source face and a target face. Diffusion models start from random noise and refine it toward a prompt; they now beat GANs (opens in new tab) on image quality benchmarks and power text-to-video systems such as Google's Veo 3, the first major-lab model to generate synchronized audio alongside video.
Large language models supply the script, producing text that is "grammatically plausible, contextually relevant and psychologically manipulative (opens in new tab)."
These tools keep getting easier to access. Criminals have moved from custom underground services toward mainstream deepfake platforms (opens in new tab) offering real-time streaming manipulation and multilingual voice cloning at low cost or free. Voice cloning can now use a short audio sample (opens in new tab) to produce a clone indistinguishable from authentic speech, with natural intonation, rhythm, pauses, and breathing noise.
A working example: an attacker pulls a short clip of an executive's voice from an earnings call or podcast, feeds it to a zero-shot cloning model, and places a live phone call in that voice, adjusting responses as the target speaks.
In January 2026, a Swiss entrepreneur lost several million francs after a series of calls in which a cloned voice impersonated (opens in new tab) a trusted business partner.
Why synthetic media matters
Synthetic media removes the cues people use to verify identity. A familiar face on a video call, a recognized voice on the phone, or a plausible document at onboarding can all appear trustworthy.
People struggle to distinguish (opens in new tab) synthetic from authentic media without technical help. In a survey of 302 cybersecurity leaders, 62% of organizations (opens in new tab) reportedly experienced a deepfake attack involving social engineering or exploitation of automated processes in the 12 months prior to mid-2025.
Financial institutions increasingly reported suspected use of deepfake media (opens in new tab) in fraud schemes by November 2024, most involving altered or fabricated identity documents built to defeat verification; per Microsoft, use of AI-generated IDs in identity fraud has grown 195% (opens in new tab) globally.
Impersonation of leaders is the enterprise use case that recurs across agency warnings. In July 2025, an unknown actor created a false Signal account (opens in new tab) impersonating the Secretary of State and used AI-generated text and voice messages to contact foreign ministers, a US governor, and a member of Congress, prompting a warning to every diplomatic post. North Korean IT-worker operatives have also used deepfake video to pass job interviews (opens in new tab) for remote roles.
Regulators now treat synthetic media as a compliance object. Under the EU AI Act's Article 50, organizations must disclose deepfake content and apply machine-readable marking to AI-generated output; most obligations apply from (opens in new tab) 2 August 2026.
Covered platforms must remove nonconsensual intimate imagery, real or AI-generated, within 48 hours of a valid request under the TAKE IT DOWN Act, which the FTC began enforcing on (opens in new tab) 19 May 2026.
Regulated financial institutions must also address AI-enabled social engineering (opens in new tab) under the NYDFS cybersecurity regulation and its October 2024 industry letter.
Types of synthetic media
Definitions differ on scope. One definition includes text (opens in new tab), while another framework restricts synthetic media to visual, auditory, or multimodal (opens in new tab) output.
The categories below follow the broader scope, since text-based personas drive many enterprise attacks:
- Deepfake video. Generation includes identity swap, attribute manipulation, expression swap, entire face synthesis (opens in new tab), and source video manipulation. People most often associate face swap with CEO impersonation on video calls.
- Voice clones and synthetic audio. Audio deepfakes cover voice cloning, speech-to-speech conversion, and text-to-speech impersonation (opens in new tab). Current systems generate speech with minimal latency (opens in new tab), so an attacker can hold a live conversation in a cloned voice.
- AI-generated and manipulated images. Face morphing blends photos of two people into one image that recognition systems match to both individuals (opens in new tab), letting one person assume the other's identity.
- AI-generated text and chat personas. LLM-generated scripts (opens in new tab) accelerate fraud, and malicious markets openly offer these tools on the surface and dark web (opens in new tab).
- Synthetic identities and fabricated documents. Attackers pair fictitious PII with an AI-generated face, then add forged documents. Injection attacks insert manipulated media directly into (opens in new tab) the digital capture stream, bypassing camera-based presentation.
- Real-time synthetic avatars on video calls. Tools such as SwapFace and Avatar AI VideoCallSpoofer fake live video (opens in new tab) calls. Platforms now treat this as a structural problem: Zoom introduced biometric "Verified Human" badges (opens in new tab) for meeting participants in April 2026.
How to defend against synthetic media
Out-of-band verification holds up regardless of how convincing an impersonation becomes. When a request for a wire transfer, payment change, or credential reset arrives on any channel, end the session and call back (opens in new tab) the requester on a number already on file, paired with a second approver and a pre-agreed code word for transfers above a set threshold.
Treat no video call as sufficient sole authorization: in one case, an automaker's executive ended a call with a cloned CEO voice (opens in new tab) by asking which book the real CEO had recently recommended, and the caller could not answer.
Training has to match the voice and video-call channels attackers use, plus messaging through messaging simulations throughout the year. Public video and audio of executives are raw material for many voice and video clones, so organizations should protect high-priority officers (opens in new tab) and monitor for unauthorized use of their likenesses.
Deepfake detection works best as one layer in a broader control stack, since an adaptive adversary can test a deepfake (opens in new tab) against a detector until it passes.
Combine detection with C2PA Content Credentials on your own published content, liveness checks paired with injection attack detection (opens in new tab) in identity verification, and phishing-resistant MFA (opens in new tab) whose token lives on a separate device from the one requesting access.
How Doppel helps
Doppel is the Frontier AI Social Engineering Defense (SED) platform that unifies Digital Risk Protection (DRP) and Human Risk Management (HRM). It treats synthetic media, deepfake video, cloned voices, and fabricated documents as one campaign: detect it externally, then train against it internally.
External impersonation campaigns spread across social platforms, messaging apps, ad networks, domains, and the dark web. Executive Protection and Brand Protection detect deepfake content, cloned voices, and profiles that reuse an executive's or brand's likeness.
The Doppel Threat Graph connects spoofed domains, fake profiles, impersonated ads, and malicious texts in a single campaign view, and its agentic AI prioritizes those signals and executes takedowns through platform APIs and escalation paths.
The same external intelligence feeds internal readiness. Doppel Simulation turns live tactics into deepfake-enabled vishing simulations and conversational scenarios across phone, collaboration, messaging, and email channels
Security Awareness Training can feature an organization's own executives from a short clip or static image, so a cloned-voice campaign detected against your CFO today can run as a company-wide simulation tomorrow. Dismantled impersonation campaigns sharpen detection for subsequent attacks, and simulations reinforce the verification reflex attackers count on breaking.
Request a demo to walk through deepfake and impersonation defense mapped to your own brand and executives.
Frequently asked questions about synthetic media
What is synthetic media?
Synthetic media is content, including images, video, audio, and text, that algorithms, usually AI, have generated or significantly altered. NIST defines "synthetic content" as information "that has been significantly altered or generated by algorithms, including by AI." The category covers fully generated output, such as a face that never existed, and manipulated output, such as a real recording with a swapped face or voice. Legitimate uses include voice banking for people losing their speech, film dubbing, and security awareness simulations.
What is synthetic media in cybersecurity?
In cybersecurity, synthetic media refers to AI-generated voices, faces, documents, and text that attackers use to impersonate trusted people and defeat identity checks. Common patterns include cloned executive voices requesting wire transfers, real-time face swaps on video calls, and fabricated identity documents submitted to digital onboarding portals. Financial institutions should flag these schemes in suspicious activity reports using the key term FIN-2024-DEEPFAKEFRAUD.
What is the difference between synthetic media and a deepfake?
Synthetic media is the umbrella term, and a deepfake is one subset. A deepfake specifically depicts a real or fictitious person, object, place, or event realistically enough to appear authentic; the EU AI Act's definition requires that the content "would falsely appear authentic or truthful," while clearly fantastical content falls outside the definition under European Commission guidance. AI-generated marketing copy, synthetic training data, and disclosed voice-banked speech are synthetic media without being deepfakes.


