Doppel Email Security is now generally available! | Register for the webinar to learn more
Research

AI-Enabled Phishing Risk Scoring: How It Works for Messages and for People

What goes into an AI-enabled phishing risk score, why message scores and employee scores answer different questions, and what each one can actually decide.

Email Threat Intelligence: From Detection to Response

The phrase "phishing risk score" covers two different measurements. One scores the message and decides whether it reaches an inbox. The other scores the employee and describes how that person has handled being targeted. Same label, same AI, no shared inputs.

A security leader who cannot say which score is on the screen cannot say what it licenses, and AI-generated phishing (opens in new tab) has raised the price of that confusion. AI-automated phishing emails (opens in new tab) reach click-through rates of 54% (opens in new tab) against 12% for standard phishing.

A workforce score read as a forecast produces false confidence, and a message score with no readable reasoning produces a block no one can defend when the business asks why a customer's invoice went to quarantine.

This article opens both scores, shows what goes into each, and marks where each stops being valid.

Key takeaways

  • Message risk scores decide delivery before anyone sees the email. Person risk scores route coaching after a program has observed behavior.
  • Scoring compresses unlike signals into one number, so a team needs the reasoning behind it and the rule that acts on it.
  • Employee scores work as coaching ledgers, not forecasts: past behavior predicts the next campaign weakly.

Two different scores wear the same name

Three things separate the two scores: what each one decides, who it acts on, and how quickly it goes out of date.

1. The message score answers whether to deliver

A message score fires before a person ever sees the email. Mail platforms map verdict tiers to actions, so a score crossing the phishing threshold routes the message to junk or quarantine, and the highest-confidence tier blocks user release entirely.

2. The person score answers who needs practice

A person score routes coaching. It summarizes what an employee did across recent simulations and real targeting events, then decides who gets the next phishing simulation (opens in new tab), on which channel, and what training follows a failure. Its horizon is weeks.

The two scores refresh on different clocks.

A message verdict describes infrastructure that often stops existing within hours. Phishing pages come down or repoint their DNS shortly after launch (opens in new tab), and attackers can weaponize links after delivery (opens in new tab). A person score moves at the pace of the program feeding it.

How a message risk score gets built

A message risk score is built from three things: signals the message itself carries, signals its sender and infrastructure carry, and a method for combining measurements that share no units. The third is where scoring systems actually differ.

1. Message signals are cheap to collect and cheap to change

Detection engines score content and urgency signals (opens in new tab) from the body and pull URL features (opens in new tab) such as character swaps like "rn" for "m." Attackers rework text and presentation in minutes, which gives these inputs the score's least durable footing.

2. Sender and infrastructure signals change more slowly

Infrastructure signals resist rewriting because they are functions of elapsed time. Domain age (opens in new tab) accrues on the calendar and sending reputation (opens in new tab) accumulates over months, so an attacker who wants an aged domain registers it well ahead of use.

Deviation from a sender's baseline requires a history the attacker did not write.

3. Compression hides which signal moved the score

A domain's age is measured in days, a link's reputation is a category, a sender's deviation is measured against its own baseline, so a scoring system has to convert all three onto a common range (opens in new tab) before adding them up.

Once it does, a benign signal can cancel (opens in new tab) a malicious one, and composite scoring inherits the problem CVSS has: many input combinations (opens in new tab) land on the same number, so the score cannot say whether the young domain or the urgent language moved it.

Attribution methods (opens in new tab) can reconstruct which feature pushed a verdict, and whether that reasoning ships with the number is the real difference between scoring systems.

How a person's risk score gets built

A person's phishing risk score is built from behavior on two paths: what someone did when a simulation tested them, and what identity and security tooling observed in production, with each event recorded against its channel.

Both paths are behavioral, so intelligence about a live campaign shapes who gets tested next rather than the record itself. Designing the program around the score belongs to awareness program guidance (opens in new tab).

Behavior observed under test

Simulations log severity-weighted events (opens in new tab). A click carries one weight, a credential submission a heavier one, and sharing an OTP heavier still, because each maps to a worse production outcome. Reporting counts in the employee's favor; reporting speed (opens in new tab) is tracked alongside it, and repeat failure (opens in new tab) compounds the record.

Credentials given up on a vishing call (opens in new tab) and a link clicked in a follow-up message (opens in new tab) are different failure modes, recorded separately.

Behavior observed in production

The second path records what happened outside test conditions, where an employee reporting a real phish earns credit for observed vigilance.

Identity and security tooling adds context simulations cannot supply, particularly when credentials surface in a breach dump (opens in new tab) or external threat data (opens in new tab) shows the workforce being targeted.

The score describes past behavior

The behavioral record has limited predictive power. Passing simulations weakly predicts (opens in new tab) future outcomes, susceptibility tracks situation more than trait (opens in new tab), and channel performance (opens in new tab) transfers poorly from one channel to another.

Read the number as a ledger of where an employee still owes practice, and against how much testing produced it: the first time multi-channel simulations (opens in new tab) run, non-email rates spike because nobody has been tested there.

What survives the compression

A score compresses reasoning into a number, and what a program can do with it depends on what survives: the threshold that converts it into an action, the reasoning a team can retrieve when the action is challenged, and how long the number stays true.

The threshold turns a number into an action

A score does nothing until someone maps ranges of it to deliver, junk, quarantine, or escalate, and that mapping is a governance decision with an owner. A missed phish stays silent until someone finds it.

A false positive interrupts the business immediately, as a quarantined invoice somebody has to clear. A team that cannot set its own thresholds has inherited someone else's tolerance.

The reasoning has to outlive the number

When the business challenges that quarantined invoice, the team needs the signals that fired: the header anomaly, the sender reputation, the URL verdict, the content pattern. An analyst holding those can release the message and narrow the policy.

Credit-scoring rules for adverse action notices (opens in new tab) require lenders to name the specific factors behind a decision, and regulators reject "failed to achieve a qualifying score" as an answer. Security deserves the same standard.

Handed a verdict that can only cite its own number, analysts ignore it or re-investigate (opens in new tab) by hand, and both responses waste the automation.

Every score has a shelf life

Attackers pre-stage infrastructure (opens in new tab) before use and iterate lures (opens in new tab) until the detection signals thin out, so neither number carries an expiry date a program can trust. Recheck a message verdict when a user acts, and attach an indicator expiry (opens in new tab) so stale ones stop firing.

Revalidate person profiles after meaningful events, not on an annual cycle.

How Doppel makes a verdict readable

Doppel, the AI-native Social Engineering Defense (opens in new tab) (SED) platform that unifies Digital Risk Protection (opens in new tab) and Human Risk Management (opens in new tab), carries both halves of this problem without pretending they are one number.

On the message side, the reasoning stays attached to the verdict, because a number that cannot explain itself forces an analyst to recreate the investigation.

On the person side, the profile records what an employee actually did, because a score earns trust only when it points back to user-level evidence.

  • Risk Modeling and Insights, inside Human Risk Management, builds a per-employee risk profile holding a personal risk score, consecutive fail streak, response speed, a per-channel breakdown, and an LLM-generated behavioral summary that recommends what to test next.
  • Employee-reported emails, real or simulated, raise a profile's standing, so the record credits vigilance as well as failure, and coaching follows the failure mode.

The two measurements stay distinct, and workflow connects them: one-click threat-to-simulation conversion turns a detected DRP threat into an employee simulation (opens in new tab) on the channel the attackers used.

Ask three questions of any score

Every vendor's score is accurate in its own demo, so the useful question is what the number carries when it reaches the person who has to act on it. Can the reasoning be retrieved? Is the threshold yours to set? Does the number still describe the world as it stands? A program that answers all three defends every action it takes on a score. One that cannot is handing its decisions to a number it cannot inspect, and as AI phishing accelerates, that gap compounds.

Request a Demo (opens in new tab) to walk through both scores against your own environment.

Frequently asked questions about AI-enabled phishing risk scoring

What signals go into AI-enabled phishing risk scoring?

It depends on which score you mean. A message-level score draws on content signals like intent and urgency, URL features like lookalike characters, and infrastructure signals like domain age and authentication. An employee-level score draws on clicks, credential submissions, and reports with time-to-report, tracked by channel.

What is the difference between a message risk score and an employee risk score?

A message risk score attaches to an email and decides whether the mail platform delivers, junks, or quarantines it, computed pre-delivery against infrastructure that can change within hours (opens in new tab). An employee risk score attaches to a person and routes security awareness training (opens in new tab) based on how they handled being targeted.

Can AI-enabled phishing risk scoring predict which employees will fall for an attack?

Person-side scoring records observed behavior, which makes it useful for deciding who needs practice on which channel and unreliable as a forecast. Susceptibility tracks situation more than trait (opens in new tab), workload predicts clicking (opens in new tab), and channel performance (opens in new tab) transfers poorly from one channel to another.

How often should a phishing risk score be recalculated?

Recheck a message score when a user acts on it. Update an employee score on event-driven triggers, not a fixed calendar.

Learn how Doppel can protect your business

Join hundreds of companies already using our platform to protect their brand and people from social engineering attacks.