Doppel Email Security is now generally available
The agentic email security solution that empowers you to fight back against social engineering attacks. Detection isn't enough. Disruption is the difference.
What goes into an AI-enabled phishing risk score, why message scores and employee scores answer different questions, and what each one can actually decide.

The phrase "phishing risk score" covers two different measurements. One scores the message and decides whether it reaches an inbox. The other scores the employee and describes how that person has handled being targeted. Same label, same AI, no shared inputs.
A security leader who cannot say which score is on the screen cannot say what it licenses, and AI-generated phishing (opens in new tab) has raised the price of that confusion. AI-automated phishing emails (opens in new tab) reach click-through rates of 54% (opens in new tab) against 12% for standard phishing.
A workforce score read as a forecast produces false confidence, and a message score with no readable reasoning produces a block no one can defend when the business asks why a customer's invoice went to quarantine.
This article opens both scores, shows what goes into each, and marks where each stops being valid.
Three things separate the two scores: what each one decides, who it acts on, and how quickly it goes out of date.
A message score fires before a person ever sees the email. Mail platforms map verdict tiers to actions, so a score crossing the phishing threshold routes the message to junk or quarantine, and the highest-confidence tier blocks user release entirely.
A person score routes coaching. It summarizes what an employee did across recent simulations and real targeting events, then decides who gets the next phishing simulation (opens in new tab), on which channel, and what training follows a failure. Its horizon is weeks.
The two scores refresh on different clocks.
A message verdict describes infrastructure that often stops existing within hours. Phishing pages come down or repoint their DNS shortly after launch (opens in new tab), and attackers can weaponize links after delivery (opens in new tab). A person score moves at the pace of the program feeding it.
A message risk score is built from three things: signals the message itself carries, signals its sender and infrastructure carry, and a method for combining measurements that share no units. The third is where scoring systems actually differ.
Detection engines score content and urgency signals (opens in new tab) from the body and pull URL features (opens in new tab) such as character swaps like "rn" for "m." Attackers rework text and presentation in minutes, which gives these inputs the score's least durable footing.
Infrastructure signals resist rewriting because they are functions of elapsed time. Domain age (opens in new tab) accrues on the calendar and sending reputation (opens in new tab) accumulates over months, so an attacker who wants an aged domain registers it well ahead of use.
Deviation from a sender's baseline requires a history the attacker did not write.
A domain's age is measured in days, a link's reputation is a category, a sender's deviation is measured against its own baseline, so a scoring system has to convert all three onto a common range (opens in new tab) before adding them up.
Once it does, a benign signal can cancel (opens in new tab) a malicious one, and composite scoring inherits the problem CVSS has: many input combinations (opens in new tab) land on the same number, so the score cannot say whether the young domain or the urgent language moved it.
Attribution methods (opens in new tab) can reconstruct which feature pushed a verdict, and whether that reasoning ships with the number is the real difference between scoring systems.
A person's phishing risk score is built from behavior on two paths: what someone did when a simulation tested them, and what identity and security tooling observed in production, with each event recorded against its channel.
Both paths are behavioral, so intelligence about a live campaign shapes who gets tested next rather than the record itself. Designing the program around the score belongs to awareness program guidance (opens in new tab).
Simulations log severity-weighted events (opens in new tab). A click carries one weight, a credential submission a heavier one, and sharing an OTP heavier still, because each maps to a worse production outcome. Reporting counts in the employee's favor; reporting speed (opens in new tab) is tracked alongside it, and repeat failure (opens in new tab) compounds the record.
Credentials given up on a vishing call (opens in new tab) and a link clicked in a follow-up message (opens in new tab) are different failure modes, recorded separately.
The second path records what happened outside test conditions, where an employee reporting a real phish earns credit for observed vigilance.
Identity and security tooling adds context simulations cannot supply, particularly when credentials surface in a breach dump (opens in new tab) or external threat data (opens in new tab) shows the workforce being targeted.
The behavioral record has limited predictive power. Passing simulations weakly predicts (opens in new tab) future outcomes, susceptibility tracks situation more than trait (opens in new tab), and channel performance (opens in new tab) transfers poorly from one channel to another.
Read the number as a ledger of where an employee still owes practice, and against how much testing produced it: the first time multi-channel simulations (opens in new tab) run, non-email rates spike because nobody has been tested there.
A score compresses reasoning into a number, and what a program can do with it depends on what survives: the threshold that converts it into an action, the reasoning a team can retrieve when the action is challenged, and how long the number stays true.
A score does nothing until someone maps ranges of it to deliver, junk, quarantine, or escalate, and that mapping is a governance decision with an owner. A missed phish stays silent until someone finds it.
A false positive interrupts the business immediately, as a quarantined invoice somebody has to clear. A team that cannot set its own thresholds has inherited someone else's tolerance.
When the business challenges that quarantined invoice, the team needs the signals that fired: the header anomaly, the sender reputation, the URL verdict, the content pattern. An analyst holding those can release the message and narrow the policy.
Credit-scoring rules for adverse action notices (opens in new tab) require lenders to name the specific factors behind a decision, and regulators reject "failed to achieve a qualifying score" as an answer. Security deserves the same standard.
Handed a verdict that can only cite its own number, analysts ignore it or re-investigate (opens in new tab) by hand, and both responses waste the automation.
Attackers pre-stage infrastructure (opens in new tab) before use and iterate lures (opens in new tab) until the detection signals thin out, so neither number carries an expiry date a program can trust. Recheck a message verdict when a user acts, and attach an indicator expiry (opens in new tab) so stale ones stop firing.
Revalidate person profiles after meaningful events, not on an annual cycle.
Doppel, the AI-native Social Engineering Defense (opens in new tab) (SED) platform that unifies Digital Risk Protection (opens in new tab) and Human Risk Management (opens in new tab), carries both halves of this problem without pretending they are one number.
On the message side, the reasoning stays attached to the verdict, because a number that cannot explain itself forces an analyst to recreate the investigation.
On the person side, the profile records what an employee actually did, because a score earns trust only when it points back to user-level evidence.
The two measurements stay distinct, and workflow connects them: one-click threat-to-simulation conversion turns a detected DRP threat into an employee simulation (opens in new tab) on the channel the attackers used.
Every vendor's score is accurate in its own demo, so the useful question is what the number carries when it reaches the person who has to act on it. Can the reasoning be retrieved? Is the threshold yours to set? Does the number still describe the world as it stands? A program that answers all three defends every action it takes on a score. One that cannot is handing its decisions to a number it cannot inspect, and as AI phishing accelerates, that gap compounds.
Request a Demo (opens in new tab) to walk through both scores against your own environment.
It depends on which score you mean. A message-level score draws on content signals like intent and urgency, URL features like lookalike characters, and infrastructure signals like domain age and authentication. An employee-level score draws on clicks, credential submissions, and reports with time-to-report, tracked by channel.
A message risk score attaches to an email and decides whether the mail platform delivers, junks, or quarantines it, computed pre-delivery against infrastructure that can change within hours (opens in new tab). An employee risk score attaches to a person and routes security awareness training (opens in new tab) based on how they handled being targeted.
Person-side scoring records observed behavior, which makes it useful for deciding who needs practice on which channel and unreliable as a forecast. Susceptibility tracks situation more than trait (opens in new tab), workload predicts clicking (opens in new tab), and channel performance (opens in new tab) transfers poorly from one channel to another.
Recheck a message score when a user acts on it. Update an employee score on event-driven triggers, not a fixed calendar.
BLOG
Detection isn't enough. Disruption is the difference. Meet Doppel Email Security: the most advanced agentic solution that progresses beyond blackbox ML and whitebox rule-based systems to detect, investigate, and disrupt social engineering campaigns end-to-end.
by Kevin Tian and Rahul Madduluri