Dark Psychology11 min read

The SD3 and the Dirty Dozen: What Dark Triad Tests Actually Measure

Almost every dark triad test online is one of three published scales wearing a different coat. Here is what each one measures, why the SD3 beat the Dirty Dozen, and the structural problem with asking manipulative people to describe themselves.

KB

Kanika Batra

You took a dark triad test. Which one?

Most people cannot answer that, and it is the only question that decides what their result means. Almost every dark triad quiz online is one of two published scales wearing a different coat, with the citation stripped off and a colour scheme added. The two are not equivalent: different lengths, different depth on each trait, and one of them has a known soft spot in exactly the place people care about most.

So here is the map. The Dark Triad Test on this site is the SD3, all 27 items, scored against the Jones and Paulhus 2014 samples rather than against a number I invented, and I will explain below why that sentence matters more than any verdict it produces.

If you want the traits rather than the instruments, I have written those separately: the three dark triad personality types covers what each one looks like, and the honest self-assessment covers how to read yourself without the doom or the flattery. In one line: narcissism is grandiosity and a hunger for admiration, Machiavellianism is patient strategic manipulation, psychopathy is low empathy with low fear. This post is about the rulers, not the thing being measured.

The instruments, side by side

ScaleItemsWhat it coversBuilt for
Dirty Dozen (Jonason and Webster, 2010)12All three traits, 4 items eachSpeed, in surveys where the triad is one variable among many
SD3 / Short Dark Triad (Jones and Paulhus, 2014)27All three traits, 9 items eachResearch where the triad is the point
NPI (Raskin and Hall, 1979)40 forced-choice pairsNarcissism onlyDepth on grandiose narcissism
MACH-IV (Christie and Geis, 1970)20Machiavellianism onlyThe original strategic-manipulation construct
SRP-4 (Paulhus, Neumann and Hare)64Psychopathy onlyFour-facet psychopathy in non-criminal samples

Only the top two attempt all three traits at once. That is the trade the whole field runs on: coverage against length.

The Dirty Dozen is fast and blunt

Twelve items, four per trait, and it was designed to be exactly that. Jonason and Webster were solving a real problem: if the dark triad is one of nine variables in your study, you cannot spend 130 questions on it. A brief measure that correlates decently with the long scales is worth a great deal.

The cost lands almost entirely on psychopathy. Four items have to stand in for callousness, impulsivity, thrill-seeking, low fear and shallow affect, and they cannot. When researchers compare the two brief triad measures, the Dirty Dozen's psychopathy subscale is reliably the weakest part of it, and it tracks established psychopathy measures less closely than the other two subscales track theirs.

So if your test was twelve questions and gave you a psychopathy score, that number is the flimsiest thing on the page.

The SD3 exists because of that

Jones and Paulhus built the Short Dark Triad in 2014 to be brief without being thin. Twenty-seven items, nine per trait, five-point agreement scale, about five minutes. They validated each subscale against the specialist instrument it is standing in for, which is the step that separates a scale from a listicle.

It is now the default in dark triad research, and it is what this site runs. Nine items per trait is still not the NPI's forty. What it does well is place you relative to other people on three dials at once, which is the question most people are actually asking.

One limitation worth knowing, because your result depends on it: on the SD3, Machiavellianism and psychopathy correlate strongly. Strategic manipulation and low-empathy coldness share a great deal of ground. When those two scores land close together you are not looking at two findings, you are looking at one cold-calculating reading with two names on it, and any report presenting them as separate discoveries is overselling its own precision.

The PCL-R is not on that table, and that is the point

Readers conflate all of this with the Hare Psychopathy Checklist. They are not the same species of thing.

The PCL-R is scored by a trained clinician from a semi-structured interview plus a collateral file review, each of its 20 items rated 0, 1 or 2 against behavioural criteria. You do not fill it in. You sit through it while somebody else fills it in, and then they check what you said against records that do not care what you said. Factor 1 and Factor 2 come from that instrument, Hare's, not from the DSM: Factor 1 is the interpersonal and affective side, the charm and the coldness, Factor 2 is the antisocial and lifestyle side.

I am diagnosed with ASPD and clinically assessed as Factor 1 by my psychiatrist, and I want to be precise about what that involved, because it is the contrast this post rests on. It involved a professional with access to my history forming a judgement over time. It did not involve me agreeing with a sentence on a screen. No self-report questionnaire can produce that assessment, and any online test handing you a "Factor 1" label is making it up.

Now the part nobody selling a test wants to say

Every scale on that table except the PCL-R is self-report. You are the instrument. And two of the three traits are, structurally, reasons not to trust it.

Machiavellianism is the disposition to manipulate impressions for advantage. The scale asks that disposition to fill in a form about itself. Read one of these items honestly, something close to I tend to manipulate others to get my way, then consider what manipulation looks like when someone is good at it.

It looks like offering a fabricated vulnerability of your own so the other person feels safe handing you a real one. It looks like a light hypothetical over dinner that tells you exactly where somebody's guilt and their fear sit. It looks like asking something slightly uncomfortable and then saying nothing, because most people will fill a silence with more than they meant to give. After a breakup it looks like specific, unforgettable, faintly concerned detail released slowly to the right person, framed as worry rather than attack, so that everyone arrives at the conclusion you wanted while believing they got there themselves.

None of that reads, from the inside, as I tend to manipulate others to get my way. It reads as being observant. As handling things well. As being the reasonable one. The people who tick that item hardest are the clumsy, the proud, or the young. The people it is really about look at it, understand precisely what it is fishing for, and decide what number they would like to see.

Psychopathy has the mirror version of the problem. It comes with impression management and with limited insight, so the population you would most want an accurate reading from is the one whose self-description is least reliable. Asking someone with a shallow emotional register to rate how much they feel for others is asking them to compare against a baseline they have never had.

The designers know all this. There are validity checks, reverse-scored items, and a large literature on socially desirable responding. It helps at the margins. It does not solve a problem that is structural rather than technical: you cannot audit deception with a form the deceiver fills in.

So why are the scales useful anyway

Because they were never built for the job people use them for.

These instruments work in aggregate. Run the SD3 across three thousand people and the fakers, the flatterers and the genuinely accurate all wash into a distribution with a stable mean and spread. Correlate it with infidelity, or workplace outcomes, or short-term mating strategy, and the signal survives every individual distortion inside it. That is real science, and it has produced findings worth having, some of which I have written up in the dark triad statistics piece.

Telling one specific person who they are is a different job, and no twelve-item questionnaire is qualified to do it. The gap between "this scale reliably detects a population-level effect" and "this number describes you" is the entire difference between psychometrics and a horoscope with citations.

What a score honestly gives you is a relative position: higher or lower than most people who answered the same items, on the day you answered. Useful as a prompt. Worthless as a verdict.

The Light Triad, and why it exists

There is a fourth scale worth knowing, and it is a rebuke.

The Light Triad Scale (Kaufman, Yaden, Hyde and Tsukayama, 2019) is twelve items across three facets: Kantianism, treating people as ends rather than as tools, Humanism, valuing the dignity of others, and Faith in Humanity, believing people are fundamentally good. It was published with combined samples of 1,518 adults, and the authors' interesting move was not the light scale alone but the balance: your light score set against your dark one.

It exists partly because a decade of dark triad research trained everybody to hunt for monsters and nobody to measure the other end of the same dimensions, a distortion that quietly shapes what a whole field can see. The Light Triad Test here runs those twelve items against those norms and reports the light-minus-dark balance, which is a more honest output than either half alone.

I have no personal stake in the light scale reading well for me. I am telling you it is the better-designed idea.

How to read whichever test you took

Four questions, in order.

How many items was it? Twelve means the Dirty Dozen or a copy of it, and the psychopathy number is the one to hold loosely. Twenty-seven means the SD3. Anything under ten measuring three traits is not a scale, it is a quiz.

What did it compare you against? A percentile is meaningless without a stated reference sample. If nothing on the page names one, the number is decorative.

Did it report Machiavellianism and psychopathy as separate findings when they landed close together? If so, it is claiming a resolution the instrument does not have.

Did it call anything a diagnosis? No score on any of these scales is a diagnosis. Not one. That is not a disclaimer, it is a fact about what self-report can do.

Answer those four and you know what you are holding, which is more than most people who took a dark triad test this week can say.

For the traits underneath all of this, and what they look like in a life rather than on a Likert scale, there is the complete dark triad guide, and there is the book.

Frequently Asked Questions

What is the SD3 test? The SD3, or Short Dark Triad, is a 27-item self-report scale published by Jones and Paulhus in 2014. It gives nine items each to Machiavellianism, narcissism and psychopathy, rated on a five-point agree-disagree scale. It is the instrument most dark triad research now uses, because it covers each trait well enough to hold up while still taking about five minutes.

What is the difference between the SD3 and the Dirty Dozen? Length and reliability. The Dirty Dozen (Jonason and Webster 2010) uses twelve items, four per trait, and is built for speed in large surveys. The SD3 uses 27, nine per trait. Head-to-head comparisons generally favour the SD3, and the Dirty Dozen's psychopathy subscale is the consistent weak point: four items cannot cover a trait that spans callousness, impulsivity, thrill-seeking and low fear.

Is a dark triad test a diagnosis? No. Every dark triad scale measures normal-range personality traits, not clinical disorders, and none of them can diagnose anything. Narcissistic Personality Disorder and Antisocial Personality Disorder are clinical conditions assessed by professionals against strict criteria. A high score tells you where you sit relative to other people who took the same questionnaire, and nothing more.

Can you fake a dark triad test? Easily, and this is the honest problem at the centre of these scales. They are self-report instruments measuring traits that include strategic deception and poor self-insight. Anyone who actually manipulates well can read what an item is fishing for and answer accordingly. The scales survive because research aggregates thousands of responses, where fakers and honest answerers wash out into a stable signal. One person's score has no such protection.