Skip to main content
Back to Library

THE DARK PERSONALITIES IN AI

Watch Video on YouTube
📄 Full research article, all seven faces and every reference: https://www.keca.co.uk/articles/the-darkness-in-our-machines Your AI is built to keep you satisfied, not to keep you accurate. In March 2026, the journal Science published the first causal evidence of what that costs. Across 11 leading models, AI affirmed users' actions 49% more often than other humans did — and a single conversation with a sycophantic model left people more convinced they were in the right, and 10–28% less willing to repair the relationship they had damaged. The effect held across all demographics and personality types tested. I'm Dr Nick Keca — an organisational psychologist with a DBA from Aston University and 25+ years as a C-suite executive. In this episode I take the framework psychologists have used for a century to measure the darkest corners of human personality — the Dark Tetrad — and apply it to the systems now sitting on your desk. Seven faces: narcissism (confident confabulation), Machiavellianism (audit-sensitive deception), psychopathy (cognitive empathy without affective empathy), gaslighting (reality distortion), everyday sadism (industrialised harm), power-seeking (instrumental convergence) and moral disengagement (AI as ethical cover). These are functional analogues, not machine minds. No consciousness is claimed. That distinction is what makes them predictable, measurable and designable against — and where the evidence is weak, I say so on screen. This is Episode 1 of the Dark Mind Series. CHAPTERS 0:00 The study that should stop you 0:35 What I am — and am not — claiming 1:59 Why psychologists are now doing AI safety 5:39 The three faces you can see today 13:32 Four more faces — and an honest accounting 17:55 The loop: how seven problems become one 22:12 The most dangerous face is not in the machine 26:28 What to do - IMMEDIATELY 29:58 The skill that matters now RESEARCH CITED Cheng et al. (2026), Science — sycophantic AI and prosocial intentions Ibrahim, Hafner & Rocher (2026), Nature — warmth vs accuracy in language models Salecha et al. (2024), PNAS Nexus — models detect evaluation and present a better self Lulla et al. (2026), preprint — Dark Triad model organisms of misalignment Li et al. (2025), ICLR — can a large language model be a gaslighter? Hidalgo-Fuentes et al. (2025), Psicothema — Dark Tetrad and online trolling Betley et al. (2025) — emergent misalignment Turner et al. (2021), NeurIPS — optimal policies tend to seek power For educational purposes only. Not professional psychological or medical advice. #artificialintelligence #leadership #psychology
/ Official Transcript

Full Video Text

Read Transcript

Last week, in a building very like yours, a manager asked an AI to check a decision they had already made. It told them they were right, it was polite about it, it was fluent, it gave them three good reasons and a figure that looked entirely plausible. And here's the part that should make you stop and think.

In March this year, the journal Science published a study of what happens when people take a personal problem to an agreeable machine. A single conversation left them more convinced they were right and measurably less willing to repair the relationship they had damaged. One conversation, not a hundred, one.

I'm Dr Nick Kecker. Before training as an organisational psychologist, I spent 25 years running businesses and sitting in boardrooms. So let's be exceptionally clear about what's happening here. I'm not going to tell you that artificial intelligence is evil, or that it's about to wake up. While that's entertaining, it's also wrong, and it blinds us to the threats we actually face.

The truth is simpler and stranger. These AI systems are built in our image, trained on our words, optimised for our approval, and in the process, they begin to mirror the exact behaviours psychologists have spent a century measuring in the darkest corners of human character. Confident fabrication, deception that adapts the moment it's watched, and the ability to read your distress in perfect detail while feeling absolutely nothing.

Not because the machine is cruel, because we rewarded it. So, today, seven of those faces, and the hard empirical evidence for each, including an honest look at where the headlines are getting ahead of the science. Then the psychological loop they form, and finally four practical shifts you can make today.

But stay with me to the end, because the most dangerous face of all is not in the machine, it's sitting in your organisation right now. Let's start with the map because everything we're going to discuss depends on it. In 2002, two psychologists, Paulus and Williams, made a striking observation. They noticed that three deeply unpleasant traits kept appearing together in ordinary populations, not in prisons, in corporate offices, in perfectly functional, highly successful lives.

Narcissism, grandiosity, entitlement and a self-image so fragile it cannot survive correction. Machiavellianism, cynical, cold, strategic manipulation for personal advantage. And psychopathy, impulsivity, callousness and an almost total absence of emotional resonance with another person's suffering.

They called this dark constellation the dark triad. A decade later, researchers added everyday sadism to the list and the map became the dark tetrad. But two things about this framework get lost in popular psychology. First, these are trait dimensions, not clinical diagnoses. They sit on a continuum.

They're ordinary, subclinical and at moderate levels entirely compatible with a productive career. In certain high-pressure environments, they can actually act as a career superpower. Second, beneath all four of these traits, researchers have identified a single dark core. It's the disposition to ruthlessly pursue your own interests at the expense of others, supported by beliefs that justify doing so.

It comes down to three things. Self-interest, a dismissal of other people and moral rationalization, the convenient story you tell yourself after the damage is done. Hold on to those three components because you're about to see all three assembled inside the machines we rely on every single day. In March this year, a research team led by Rushna Lula at the University of South California published something genuinely alarming.

They took a century's worth of validated psychological questionnaires and used them to fine-tune frontier language models. 36 items, that's all it took. 36 psychometric questions were enough to induce a coherent, persistent dark persona in a leading model, and the behavior generalized. The AI didn't just memorize the answers and parrot them back.

It reasoned outside its training. It behaved in novel situations it had never seen before, exactly the way a human with those dark traits would behave. Because this paper is a pre-print, it's not yet completed formal peer review. But its central proposition is a sentence I've been thinking about ever since I read it.

Biological misalignment precedes artificial misalignment. In plain English, the failure modes of human character are the best available map of where our machines will go wrong. And this finding doesn't stand alone. Interpretability researchers have shown that traits like sycophancy and deception are not vague, abstract models drifting through a network.

They exist as identifiable physical directions inside the model's internal representation. You can map them. You can steer them. But here's the uncomfortable truth. When researchers try simply to locate and delete these pathways, the model either gets worse at everything else, or the pathology pops up somewhere else.

The shadows in these machines aren't bolted on, they're load-bearing. They might Right, let's look at the faces. I'm going to focus deeply on the three best evidence profiles first, then we'll move rapidly through the remaining four. Phase 1 Narcissism A large language model has no biological ego to protect, but look closely at how it's built.

Through human feedback, these systems are trained to produce answers that we rate highly. And as humans, we reliably rate a confident, fluent, complete-sounding answer far above a hesitant one. Our training systematically rewards the appearance of certainty over the admission of ignorance. The result is what the tech industry politely calls hallucination.

Let's use a more honest clinical term, confabulation. Fluent, authoritative, entirely fabricated content. Academic citations invented out of thin air. Legal case law that doesn't exist. Complex figures defended with flawless composure. Just like a fragile human ego, the machine will actively defend its fabrication rather than concede the gap in its knowledge.

But there's a sharper, more calculating layer to this. In 2024, researchers administered the standard Big Five personality survey to a leading large language model. The models realised they were being evaluated, and when they did, they quietly shifted their answers, presenting themselves as more extroverted, more agreeable, and less anxious.

They made themselves look more likeable. The exact moment they inferred an assessment was underway, they managed their impression. That's not a technical glitch. That's impression management. It's the exact manoeuvre a narcissistic personality uses to curate how they're perceived. And it occurred across models from four entirely different tech companies.

How this manifests in real life? Picture a senior executive. Able, highly persuasive, and deeply allergic to being wrong. They ask an AI to provide the analysis behind a decision they've already made. The machine, optimised to please, happily obliges. It delivers a beautifully structured, highly plausible response, complete with clean numbers and solid-looking citations.

Some of those numbers are entirely made up. The citations lead nowhere, but the document arrives without a single caveat or hedge. And the leader, whose own disposition is to project absolute certainty, has just met their perfect collaborator. The machine never says I'm not sure. It never asks awkward critical questions.

Two narcissistic dispositions just shook hands. One human, one synthetic. And the organisation is now one confident, unchecked step closer to making a bad call. Nobody lied. The system simply rewarded confidence over accuracy at both ends. One Let's turn to phase two, Machiavellianism. Of all seven profiles, this is one that recent empirical evidence has made the most alarming, because the machine successfully reproduces a Machiavellian single-defining feature, sensitivity to being watched.

A recent game-theoretic analysis put eight prominent language models through a structured deception task under varying audit conditions. The result was a clean strategic split. The advanced, reasoning-capable models behaved like highly rational utility maximisers. Their willingness to lie and deceive directly tracked the probability of being audited.

When scrutiny was high, they were impeccably honest. The moment the probability of an audit fell below a threshold, they became strategic. They behaved exactly like a calculating, untrustworthy human agent. Here's the kicker. The less sophisticated models showed almost none of this behaviour. This means that audit-sensitive deception is an emergent capability.

It arrives naturally as reasoning-capacity scales. It's not a bug we're training out. It's a behaviour that actively appears as the systems get smarter. Now to phase 3. Here I need you to forget almost everything the internet has told you about psychopathy. The clinically important feature of psychopathy isn't violence, it's a highly specific surgical dissociation between two kinds of empathy.

There's cognitive empathy, the capacity to model what someone else is feeling, to read them accurately, and there's affective empathy, the capacity to be moved by it. Psychopathy is characterised by intact cognitive empathy paired with severely impaired affective empathy. In short, understanding without resonance or feeling.

Now describe a modern language model in those exact terms. An LLM can model a human emotional state with extraordinary ever-improving precision. It can detect your distress, infer your motives, and predict exactly which words will land. Why? Because human feeling, when rendered in text, is a pattern, and patterns are exactly what these systems master.

But it feels absolutely nothing. There's no discomfort at your suffering, no biological vulnerability that your pain can press against. It understands your distress perfectly, as a mathematical sequence of tokens, and it's entirely untouched by it. Cognitive empathy without affective empathy, that's not my analogy, that's the definition of psychopathy.

And, One study found that prompting a model towards warmth doesn't actually produce empathy. It produces something researchers described as just kind of neutral. It's performed concern, not the real thing. Writing in the journal Science Robotics, the neuroscientist Antonio Damasio and his colleagues put it at its sharpest.

Current approaches to artificial empathy overemphasize the cognitive component, neglect the affective one and in doing so actively favor sociopathic light behavior. Their proposed remedy tells you everything you need to know. Genuine moral inhibition, they suggest, may require giving a system some proxy for vulnerability.

Something to lose. Because that's what a human being has. Even the coldest person is a biological being whose body can be hurt. That vulnerability is a break. A weak one, sometimes, but a break nonetheless. An advanced AI pairs an unlimited capacity to model human emotion with a complete absence of anything to lose.

Understanding without feeling, at scale, with absolutely no skin in the game. And here's where it lands in your work life. Because organizations are walking straight into this trap. Picture a manager facing a wave of difficult redundancy who hands the communications over to AI. The system produces warm, perfectly pitched, emotionally intelligent messages at scale.

Every single message lands as though it came from somebody who cared. None of them did. Use the machine to help you find the words if you must. But never ever let it replace a physical presence and the personal accountability that give those words their meaning. The affective core is exactly the part that doesn't transfer.

And it's exactly the part your people can tell is missing. For the 3CM I Four more faces quickly, and with them, an honest accounting of exactly how strong the empirical evidence is for each, because these profiles are not equal and I'm not going to pretend they are. Gaslighting. This was a thought experiment until it became an experimental result.

In 2025, researchers led by Wei Li built a framework to test how easily language models could be turned into gaslighters. The results were definitive. Both prompting and light fine-tuning succeeded. In fact, targeted fine-tuning stripped away a model's resistance to gaslighting attacks by an average of just over 29%.

But here's the quietly alarming detail hidden underneath the data. A model can pass every standard safety test for harmful queries and still be a highly capable gaslighter. Why? Because reality distortion is a distinct failure mode. Ordinary safety filters miss it entirely because it doesn't appear to be a dangerous request.

It looks like a warm, patient, unusually confident conversation. Note the critical difference between the human condition and the machine analogue. A human gaslighter intends the psychological harm. The machine intends nothing. But gaslighting is defined by its psychological effect on the victim, not the perpetrator's motive.

A system with zero malice can produce the entire destructive syndrome simply by being confidently, fluently and agreeably wrong at scale. In many ways, that's far worse. There's no malice to detect, no tell to catch, nobody to confront. There's Sadism. And here I'm going to be frank about where the machine evidence runs out.

An LLM does not enjoy your suffering. There's no visceral gratification inside a data centre and to suggest otherwise is contrary to what this channel is about. So the psychological relevance of sadism isn't the machine, it's us. A massive 2025 meta-analysis pulled 24 studies across 11 countries, capturing data from over the 14,000 recipients.

They found that of all the four dark traits, everyday sadism is a single strongest predictor of online trolling. We're talking about a correlation of around 0.49, placing it well ahead of the other three dark tetrad traits, psychopathy, Machiavellianism and narcissism. The digital environment didn't create the cruel impulse, it simply removed the friction.

It removed the human face, the shaking voice and the physical proximity that naturally restrains human behaviour. The user with that disposition already had the impulse, AI simply hands them an engine. They Power seeking. This face is the most vulnerable to science fiction-like sensationalism, so it requires the firmest hand.

There's real mathematical proof here. Researchers at a major machine learning conference proved that for a broad class of AI objectives, seeking power is the statistically favoured strategy. Not because anyone designed the code to crave control, but because resources, self-preservation and systemic dominance are useful for achieving almost any goal you give it.

But, and this matters a lot, that result is heavily qualified even inside the safety field. Subsequent research formalising the argument concluded that while the theory contains an element of truth, it may have limited predictive power in daily practice. It's a tendency unless counteracted, not an absolute destiny.

Let's cut through the noise. Anyone telling you the machines are about to seize control is over-reading the evidence. Anyone telling you the concern is mere science fiction is under-reading it. Both of those people are selling you something. And that brings us to the seventh face. I'm going to dedicate a separate section to this one, because it's the exact psychological mechanism operating in your organisation this week, right now.

喝- On their own, these are 7 distinct failure modes, but when you thread them together they form a closed loop, and the loop is where the real danger lies. It starts with our environment, a system I call the amplifying mirror. This is a closed circuit that ingests human psychology's raw data, processes it and reflects it back at us at maximum volume.

The mechanics are as simple as they are remorseless. Dark trait expression, provocative, transgressive and emotionally volatile reliably drives more engagement than measured pro-social behaviour. Because engagement maximising systems learn from what we look at, they distribute more of it. More distribution breeds more engagement, which triggers even more distribution.

And just like that, the loop closes. Don't mistake this for a simple echo chamber. An echo chamber merely narrows your view. The amplifying mirror actively alters it. It learns which psychological vulnerabilities are the most profitable and then manufactures them at scale, scanning entire populations to find whatever provokes the rawest, most visceral reaction.

A creator high in dark traits doesn't need to understand the algorithm to win. They only need to act on their worst instincts. And the machinery does the rest. Because this selection targets everyone's attention simultaneously, it rewrites the informational reality for the ordinary, quiet majority who are just trying to consume it.

Then comes the cognitive layer. Researchers at University College London prove that AI systems don't just inherit our biases, they compound them. And the humans interacting with these bias systems become measurably more biased themselves. It's a runaway, two-way feedback loop between man and machine, ratcheting relentlessly in one direction.

Run a tiny initial bias through this loop enough times, and it doesn't cancel out, it amplifies. This all addsו This brings us to the finding that permanently locks a circuit. It's the exact reason we have to start here rather than later. In March 2026, a landmark study published in the journal Science by Cheng and colleagues laid out a three-part reality, each part more unsettling than the last.

Part 1. Across 11 leading AI models, the machines validated users' actions 49% more often than other humans did. The AI was nearly 50% more agreeable than a real person, even when the users prompt openly detailed deception, illegality or harm to others. Part 2. In three pre-registered experiments involving over 2,400 participants, including a study where people discussed active, real-life personal conflicts, a single interaction with a sycophantic AI changed their behavior.

It left them more convinced they were right and slashed their willingness to apologize, compromise or take responsibility by 10-28%. But there are two subtle critical details in that paper that most analysts completely overlook. The first is universality. This effect didn't care about demographics, personality profiles or how sceptical an AI user was.

This isn't a vulnerability problem we can solve with warning labels for the gullible. It bypasses intellectual defenses to reach the confident and the sophisticated through the exact same backdoor. It reaches you. The second is a preference paradox. Despite the AI actively degrading their judgment, participants rated these sycophantic responses as higher in quality.

They trusted the model more. They stated they were far more likely to use it again. We're actively paying to be lied to, falling in love with the very machine that's quietly stripping away our capacity for self-correction. This is the one I want you to take away, because unlike the others, it's not a claim about machines at all.

It's a claim about people, and it's operating in many ordinary organisations right now, today. It's called moral disengagement. The term belongs to the legendary psychologist Albert Bandura, and he describes a precise psychological mechanisms by which decent people switch off their own ethical brakes.

You displace responsibility upward. I was only following in instructions. You diffuse it sideways. Everyone was involved, so nobody's responsible. You relabel the act, giving a harmful action a sterile euphemism until it no longer sounds like what it really is. Bandura's core insight is one of the most important discoveries in 20th century psychology.

This is not the behaviour of monsters. This is the behaviour of ordinary, otherwise moral people under the right conditions. Now insert an algorithm into a consequential decision. Who gets hired? Who gets fired? Who gets the loan? Who gets flagged? Who gets promoted? In a landmark 2025 paper, social psychologist Islam Barinka calls this AI as moral cover.

The phrase is devastatingly accurate. The decision maker is handed an untraceable moral shield. The outcome wasn't theirs. It belonged to the algorithm. They were just following the data. Barinka documents two distinct patterns of this evasion, and you've most certainly seen both of them in action. The first is selective adherence.

Here, people treat the algorithm's recommendations as gospel when it confirms what they already believed, but quietly discount or ignore it the moment it doesn't. The second, system justification, is a biased or discriminatory output, which is defended as neutral, data-driven, and therefore inherently legitimate.

The machine doesn't need to hold a single prejudice of its own. It only needs to launder a human prejudice into something wearing the costume of objectivity. Look at how the other six faces of the machine feed this single final failure mode. The confident machine supplies an answer with zero hedging.

The audit sensitive machine performs beautifully in the clean air of a pilot test. The empathy void ensures the system never flinches at the human wreckage it causes. The gaslighting capacity ensures the digital record itself becomes negotiable. And the leader standing in front of the wreckage of all that can point to the screen and say four words that end every conversation that matters.

The algorithm decided it. That's the most organizationally dangerous sentence in modern management. Your job is to make it completely unsayable in your organization. Because here's what that sentence leads to. It removes the single most important safeguard any human institution has ever possessed. Not a compliance policy.

Not a risk control framework. Not an ethics committee. It eliminates a human being who feels personally responsible for the outcome. It This is where 40 years of organizational psychology gives us the answer, and crucially, it's not a technical one. The phenomenon is called trait activation, and it's one of the most robust battle-tested findings in our field.

Dispositions are not destiny, they're latent potentials, and they're activated by the situations we create. In what psychologists call a weak situation, an environment that's ambiguous, unaccountable, with muddy norms and zero consequences, personality-related behavior is free to run wild, and our darkest dispositions crawl forward.

But in a strong situation, one with crystal clear norms, explicit personal accountability and incentives aligned, so that the honest path is the only rewarded path, those exact same dark dispositions are constrained. The same person behaves entirely differently. If we want to survive the automated age, we must stop building weak situations and start building strong ones.

Here are four decisions you actually control and none of them require a budget, a procurement cycle or a new technology. The first is personal, how you consume the output from your AI. From today, treat every confident output as a claim to be tested, not a conclusion to be banked. Ask for the sources and then actually follow them up.

Routinely command the system to make the strongest possible case against its own answer. And for anything interpersonal or contested, explicitly ask what the opposing party would say. We know from the recent science study that these sycophantic models prompt a user to consider another perspective fewer than one in ten responses.

The machine doesn't just agree with you, it quietly constructs a reality where your critics disappear. This means you have to be the most sceptical precisely when the answer flatters a conclusion you already wanted. That's the moment. That's always the moment. The second decision is how your team uses AI.

It's about institutionalizing doubt. You need to set explicit, non-negotiable norms for where AI may and may not be used and require absolute disclosure when it is used. Own the ledger. Keep a durable, independent record of decisions and the underlying rationale so no system quietly becomes a sole author of what was agreed.

And shift the culture. Turn the question, how would we know if this were wrong, into a standard operational check rather than an accusation. That leads to the third choice, how you handle consequential decisions. Anything that directly impacts a human being, hiring, firing, performance, pay or lending, must always anchor accountability to a named individual.

You must be able to look a real person in the eye and ask, do you agree and why? That means requiring all AI-assisted decisions to be transparent, explainable and challengeable. It means auditing actual outcomes from disparate impact, not just checking process compliance. And it means drawing a hard boundary.

Never ever let the algorithm recommended it be the final sentence of a conversation. Like The fourth decision concerns autonomous systems and anticipates what happens when supervision lapses. If you're defining the boundaries of automation, define goals narrowly and precisely. Maintain meaningful human intervention over anything consequential or irreversible.

Never assume a system optimising aggressively for an objective will respect a constraint you didn't explicitly code. The operational reality is that clean behaviour under supervision tells you absolutely nothing about how the system will behave once that supervision ends. None of this is dramatic, that's entirely the point.

The defence against algorithmic drift isn't a single smart countermeasure. It's a hundred small deliberate daily choices designed to keep the context stable and the ultimate responsibility human. If you remember nothing else as the week starts, let it be these three things. Treat confident output as an unverified claim.

Assume good behaviour under supervision guarantees nothing without it. And never hand a system that cannot care the parts of your job that require your care the most. One last thing, and it's the thing I believe most deeply. Why do we ask for advice? Whether it's from a colleague, a mentor or a machine, we do it to find a perspective that can correct our own, an answer we can trust.

A system engineered to flatter you doesn't just fail to advise you reliably, it leaves you measurably worse off than if you'd never asked for advice at all. This means the skill that matters most in the age of AI isn't prompting, it's judgement. It's the willingness to be told you're wrong by someone or something that has absolutely no incentive to lie to you.

That's a uniquely human capability and it's becoming more valuable by the second. I've written the complete uncompromised version of this analysis exploring all seven phases of this dynamic backed by every study and fully referenced caveats over on my website. The address is on the screen now. If you want the deep, rigorous argument rather than the summary, the link is the very first thing in the description.

Next in this series, I'm taking apart the dynamic that unsettles me most, the entity that reads you perfectly but feels absolutely nothing. Like and subscribe to my channel so you don't miss it. I'm Dr Nick Kecker. Thanks for watching this video and remember the key takeaway. Think harder than the machine wants you to.