Abstract
The question this article answers is: when somebody or something describes what you are like, what is that description actually worth, and what does it entitle anybody to do?
Observer reports, meaning descriptions of a person’s character given by somebody other than that person, are among the best-evidenced and least-used instruments in personality psychology. This article appraises what they establish about the subclinical dark traits, then follows the same inferential act into two places the original research did not anticipate: the one-dimensional commercial personality report, and automated inference from digital traces.
Its evidential spine is a 2025 pooling of 24 studies and 7,022 people, the first to bring together every study assessing the dark triad through observer reports, which finds convergence of medium to medium-high size for all three traits, strongest for psychopathy and weakest for Machiavellianism, with acquaintance changing the picture for two of the three and not for narcissism (Rico-Bordera, Pineda, Galán & Piqueras, 2025). That is placed against the wider literature on judgement accuracy, in which intimacy rather than frequency of contact carries accuracy, observer descriptions predict job performance more strongly than self-descriptions, and the addition runs in one direction only (Connelly & Ones, 2010; Oh, Wang & Mount, 2011; Ashton & Lee, 2025).
The article’s central structural claim is that employers borrowed a science built on domestic relationships, applied it to the relationship it identifies as weakest, and reduced it in transit to a format that omits what makes a trait score interpretable. Trait activation and trait interaction are treated as mechanism rather than taxonomy, through the leadership case in which the association rises and then declines, and through the live dispute over whether the dark traits are largely a region of ordinary personality rather than a separate territory.
It then examines algorithmic inference on the same terms rather than as an ethical appendix: smartphone traces (Marengo, Elhai & Montag, 2023), social media likes and language (Youyou, Kosinski & Stillwell, 2015; Peters & Matz, 2024), dark-trait models built from posts (Leberecht et al., 2026) and automated video interviews (Hickman, Bosch, Ng, Saef, Tay & Woo, 2022). The comparison is not which observer wins. It is that a model has less access to circumstance than a colleague, returns a score of exactly the one-dimensional kind this article criticises, and cannot be asked what it saw.
A dedicated late chapter sets out who is wrongly convicted on this material. The practical section is written for both sides, because most readers are judged as often as they judge, and it closes on entitlement rather than accuracy: what a description permits, who may collect it, and whether the person described can answer it.
Keywords: dark triad; observer reports; self-other agreement; trait activation; digital phenotyping; algorithmic personality inference; workplace monitoring; organisational psychology.
PART I. WHAT A PERSON CAN SEE
Chapter 1: The Sentence at the Kitchen Table
1.1 Sunday
A note on what follows. The person in this chapter is a composite. The sequence combines patterns I’ve seen repeatedly in my career, and no identifying detail belongs to any individual. She is here because the evidence needs a human problem to keep complicating, and not because her case proves anything.
I’ll call her Mary, a woman of thirty-eight; she’s just been appointed to run a product team after a fast rise, and on Sunday evening she is at a friend’s kitchen table because a third friendship has ended the same way the previous two did: in a silence nobody explains. The friend says the sentence she has been holding for about two years. This keeps happening to you. Mary hears an accusation, and the friend believes she is finally being honest about a pattern; both are partly right, which is why the evening doesn’t recover.
1.2 Monday
On Monday, the appointment is confirmed. Attached to the internal file are two references describing her as decisive and commercially strong, an appraisal written by one manager, and an automated assessment generated from her recorded interview, which describes her as high in drive and low in interpersonal caution.
By breakfast, five descriptions of Mary exist, and they do not agree in any respect that matters. Hers is that she is direct and calm under pressure, her friend’s is about a pattern running across fifteen years, and a former partner’s used the word controlling for behaviour Mary experienced at the time as nothing more than clarity.
Her referees’ descriptions are favourable and were chosen by her, which is the condition under which references are always collected. The fifth was produced by a system, based on the timing and language of her answers, and nobody in the process can explain how; she has not seen it.
Only one of those five was generated without anybody needing to know her at all. This article is built around that fact, and it’s a new and emerging situation we are all facing.
1.3 What the Five Have in Common
Each is a claim about what Mary is, based on something she did. Each was made by an ‘observer’ with particular access and a specific interest. And each will be weighted by whoever is deciding, mostly on the basis of where it came from rather than how it was produced.
The friend holds fifteen years of observed behaviour across a dozen different contexts and carries no professional standing. The referees have professional standing, limited closeness, and were selected by the person they are describing. The system holds more raw observation than the other three put together and no understanding of a single circumstance that produced any of it.
Ask which of them knows her, and the question doesn’t resolve. Ask instead what each of them can actually see, what each of them wants, and whether Mary can answer any of them, and the picture becomes tractable. Those three questions run through the rest of this article.
Chapter 2: How Much of Somebody Shows
2.1 The Question Nobody Had Pooled
Narcissism, Machiavellianism and psychopathy, in the everyday subclinical sense rather than anything a clinician would diagnose, are not neutral descriptions. They carry a negative stigma, and the standing objection to measuring them with questionnaires is that anyone with a reason to conceal them will simply answer as they wish to be seen. This is known as impression management. The obvious answer is to ask somebody else, and until recently nobody had pooled the studies that did so. That gap was closed in 2025, when a team at Miguel Hernández University searched four databases for every study measuring any of these traits, both by asking the person and by asking somebody who knew them, found 24 that qualified, covering 7,022 people, and combined the results (Rico-Bordera, Pineda, Galán & Piqueras, 2025).
Their stated purpose was to test whether observer reports could supplement self-reports or replace them, and that is a strong question to ask about a method most organisations treat as hearsay.
2.2 What They Found
All three traits converged between the two viewpoints, at a level personality research treats as substantial, with the callous and impulsive cluster highest, narcissism next and manipulativeness lowest. The exact figures are in the evidence note at the end rather than here, because a number nobody repeats is not a finding a reader can use.
The comparison is the useful part. The equivalent convergence for ordinary personality traits, the ones nobody has any reason to hide, typically runs lower than what these three produced. Some socially undesirable traits seem more visible than most people expect, which is close to the reverse of the popular assumption.
The authors explain that a characteristic becomes easier to judge when it is strongly desirable or strongly undesirable, because both attract attention and get talked about. You can check that against your own memory. Nobody in an organisation recalls who was reliable in March, and everybody recalls who took the credit for something in front of a room.
2.3 The One the Room Reads Worst
The three did not come out equally, and the order is not the one most people would guess. Manipulativeness produced the weakest convergence of the three, and callousness the strongest. The popular treatment ranks these traits by how frightening they sound, which puts psychopathy first and manipulativeness behind it as the lesser problem. The evidence on visibility ranks them by how much of the person reaches others, which puts the strategic one at the far end, where the light does not reach.
The authors’ account matters because it changes what you look for. Somebody high in narcissism has a reputation to maintain and behaves consistently across audiences, which makes the behaviour legible. Somebody high in psychopathy is not much concerned with how they come across, so the behaviour is not managed either. Somebody high in manipulativeness is doing something else entirely, because the trait is characterised by strategic self-interest, so the behaviour is shaped for whoever is watching, and every individual witness receives an accurate account of a different presentation.
There is one further asymmetry. For manipulativeness and for the callous cluster, closer raters agreed more with the person’s own account. For narcissism, how well the rater knew the person made no reliable difference in either model the authors ran. That does not make a brief acquaintance a generally good judge of character; it means that on that one trait, distance costs less than it does elsewhere.
2.4 What This Does Not Show
Here the article has to be careful, because the finding is easy to over-read and the over-reading is the version that travels. Agreement is not accuracy, and what these studies establish is that two viewpoints converge, and the thing each observer is compared against is the person’s own questionnaire score, which is the very instrument everybody agreed was suspect at the start of this chapter. If a person and the people around them are wrong in the same direction, the method records agreement and reports a success.
The pooling carries limits its authors name. Twenty-four studies is a small pool, none of the designs followed anybody over time, and only one study covered sadism, so nothing can be said about the fourth trait. And two of the twenty-four used people at work, which becomes the subject of Part II.
Chapter 3: What Makes a Witness Good
3.1 Closeness, Not Exposure
If you had to name the person best placed to describe a colleague, you would probably name whoever spends most time with them. That answer contains an assumption which has been tested and found wanting.
Three pooled analyses covering 263 independent samples and 44,178 people separated two things that ordinarily travel together: how often you interact with somebody and how close you are to them (Connelly & Ones, 2010). Interaction frequency improves accuracy modestly, while interpersonal intimacy produces substantial gains.
Closeness does not help evenly. It helps most for traits that are hard to see from outside, where you describe how somebody experiences the world rather than what they do in it, and least for traits that amount to a verdict rather than a description. The three traits this series covers sit squarely at the evaluative end, which places a ceiling on what closeness can buy you in exactly the territory where a reader most wants it to help.
Then the finding that no organisational practice reflects. Family and friends produce the most accurate descriptions, and colleagues, despite far more frequent contact, produce weaker ones. The workplace is a setting of high frequency and low intimacy, which is precisely the combination identified as producing plenty of observation and modest accuracy.
3.2 The Friend Who Has Never Been to Your Workplace
The obvious objection is that somebody outside work cannot possibly say anything useful about behaviour inside it. That objection has been tested directly, and it does not survive.
A German study asked 111 employees to rate themselves, then collected ratings from 106 personal acquaintances, including family members, and from 102 co-workers. Both kinds of outside observers produced comparable descriptions of worth, suggesting that an observer's advantage is not contextual knowledge of the job but a less obstructed view of the person (Connelly & Hülsheger, 2012). The same study found that people who overestimated their own agreeableness and conscientiousness performed worse than those who did not.
Two later studies pushed it further, comparing acquaintance ratings of personality against workplace misbehaviour as recorded by the person’s own supervisor. The acquaintance ratings predicted that misbehaviour, and they did so over and above what the person said about themselves (Kluemper, McLarty & Bing, 2015). What your friend lacks is any knowledge of your job. What they have is an accurate picture of you, which is the harder half to obtain and the half your employer is not collecting.
3.3 Who Is Doing the Describing
Another variable almost nobody controls for is the observer themselves. In a field study of 301 people, ratings of a target’s personality predicted that target’s supervisor-rated job performance more strongly when the observer was higher in conscientiousness, openness and emotional stability, while the observer’s own agreeableness and extraversion made little difference (Klinger & Siangchokyoo, 2024). Two people can watch the same colleague for the same three years and produce descriptions of measurably different worth.
That is uncomfortable in a way the earlier findings are not, because it means the quality of an account depends partly on a property of the accuser that nobody in the room is assessing. It also suggests a practical filter that costs nothing: the careful, curious, steady colleague is a better witness than the one who is most certain.
3.4 A Third of Any Account Is the Person Giving It
Everything so far argues for taking other people’s accounts more seriously. This is the counterweight, and without it the argument becomes dangerous.
When several people describe the same person, what they say separates into parts. Some is shared by the person and everybody rating them, which is the closest thing available to the trait itself. Some is what the person believes, and nobody else sees. Some is what observers see, but the person does not, which is the reputation (McAbee & Connelly, 2016; Connelly & McAbee, 2024).
Recent work quantified the remainder. Across two samples carrying 1,615 and 1,434 informants, a substantial share of what any single rater said about somebody was unique to that rater, shared neither with the person nor with the other raters, running for most facets at between a quarter and a third (Wiedenroth, Connelly, McAbee & Fang, 2025).
That does not license the arithmetic that two accounts are worth twice one, and I made that error in an earlier draft of this article. What it licenses is smaller and still useful. A single account always carries the person giving it, in a proportion large enough to matter, and accounts that share an origin are closer to one account than to several.
There is a further reason to distrust a single source. When people know their description of somebody will be read or used, their account moves in a favourable direction, an effect demonstrated experimentally and named after the letter of recommendation (Leising, Erbs & Fritz, 2010). The pause at the kitchen table is worth more than the written reference for exactly that reason, because nothing was going to be done with it.
PART II. HOW THE SCIENCE LEFT THE ROOM IT WAS BUILT IN
Chapter 4: A Domestic Science Borrowed by Employers
4.1 Where the Evidence Actually Comes From
Read the studies behind Part I, and the striking aspect is who did the rating. Parents, friends, romantic partners, roommates, classmates and, in a couple of cases, prison officers. Only two of the twenty-four studies in the dark-trait pool used people at work.
This is primarily a science of domestic life. It was built to answer how people who share a home, a history, or a friendship come to know one another, and its central finding is that knowing depends on sharing.
What happened next is the structural problem at the heart of this article. Employers borrowed the findings almost exclusively and applied them to the one relationship the research identifies as producing the weakest readings. Then they quietly used the domestic evidence to justify the workplace application.
4.2 What Got Borrowed and What Got Left
Three things travelled from the research into practice, and they were the three easiest to sell. The idea that personality predicts performance travelled, because it justifies a purchase. The questionnaire travelled, because it scales. And the trait names travelled, because they are memorable and they make a report readable.
What did not travel is everything that makes the first three interpretable. The dependence of accuracy on closeness did not travel, which is why the reference call reaches a former manager rather than anybody who knows the candidate. The one-directional nature of the evidence did not travel: adding other people’s descriptions to a self-description improves the prediction of performance, and adding the self-description to the others’ does not (Oh, Wang & Mount, 2011; Ashton & Lee, 2025). Every selection process I have seen collects the expendable half of that pair first and most carefully.
4.3 The One-Dimensional Report
Here is what the borrowing produced, in the form most readers will have met it. You complete a questionnaire. You receive a set of scores. Somebody explains what the scores mean, usually in a page of prose that reads as though it were written about you and was written about a band.
That format is not a simplification of the science. It is a different claim. The questionnaire produces a position relative to other people, and the report converts it into a description of a person, which is a move the underlying research does not support and which the next chapter is entirely about.
It is also where the access problem lives. Interpreting a trait score properly requires knowing how traits behave in situations, how they interact with each other, and where the relationship between a score and an outcome stops being a straight line. That is specialist knowledge; most people do not have it, and the report does not supply it. So a reader is handed a number and left to do the one thing the number cannot support: conclude something about themselves or somebody else as an individual.
Chapter 5: Why a Score Cannot Tell You About a Person
5.1 The Man Who Was Impossible Until He Was Not
A finance director I worked with had been, by common account, impossible for four years, and by the time I met the team, the description had hardened into the kind used in his absence and never in his presence. He interrupted, took over meetings that were not his, and made a habit of arriving with a version of a decision already settled so the discussion became a formality. He moved to a subsidiary with a different chief executive and a different set of rewards, and within a year the same man was described by the people around him as demanding but straight, a phrase nobody in the previous business would have recognised.
Nothing about him had changed, and there is no version of a trait score that could have told you which of those two years you were about to get. What changed was that the first business paid for certainty and the second paid for being right, and a tendency needs an occasion before it becomes behaviour anybody can see.
5.2 What the Report Does Not Say
Now take that man and imagine the document. He completes a questionnaire, receives a set of scores, and somebody sends a page explaining what the scores mean, written in prose that reads as though it were about him and was, in fact, written about a band.
The page cannot tell you that the behaviour needs a reward before it appears, because the questionnaire did not ask about his employer. It cannot tell you that the same score in the subsidiary would have produced a different year. And it will be read, by him and by everybody else who sees it, as a description of the man rather than as an estimate of what he tends to do when the circumstances invite him.
This is where the access problem lives, and it is not that people are incapable of understanding personality research. It is that the interpretation requires knowing what activates a trait, how traits combine, and where a relationship between a score and an outcome stops behaving in a straight line, and the report supplies none of the three while implying that none is needed.
5.3 The Score That Means Two Opposite Things
The clearest case is leadership. Pooling the research on narcissism and leadership, the association does not run in a straight line: it rises as the trait rises, reaches a point and then falls away, so that moderate scorers are rated as better leaders than either the low or the high scorers (Grijalva, Harms, Newman, Gaddis & Fraley, 2015). Read that against the page of prose, and the problem is not subtlety; it is direction. A bar three-quarters along a scale sits on the downward slope or the upward one depending on where the turn falls; the document does not mark the turn, and two people with the same number are therefore not the same case in any respect that matters to whoever is about to promote one of them.
I have never seen a commercial report that shows the turn. I have seen many that describe a high score as a strength and a very high score as a strength, expressed enthusiastically.
5.4 Whether These Traits Are Their Own Thing at All
A disagreement underlies all of this, and it belongs in the open, because how it resolves changes what a reader should believe. One camp holds that what the three traits share is essentially the opposite pole of agreeableness, so that the dark triad adds little beyond a trait every ordinary personality model already measures, and that the apparent distinctiveness of the dark items is largely a matter of how they are worded (Vize, Miller & Lynam, 2021). The other holds that there is a common dark core in its own right, related to low honesty-humility but not reducible to any single ordinary trait, and that this core, rather than agreeableness, explains what the three have in common (Moshagen, Hilbig & Zettler, 2018; Moshagen, Zettler, Horsten & Hilbig, 2020).
If the first camp is right, the popular treatment of these traits as a separate species of person is wrong at the root, and what the reader is actually looking at is the low end of a dimension that everybody sits on somewhere. This article treats them as useful descriptions of behaviour rather than as categories of person, which is the reading both camps permit and the only one that survives either outcome.
5.5 Population-Grade Instruments, Person-Grade Decisions
Put the finance director, the page of prose and the turn in the curve together, and the charge becomes specific. These instruments work on groups, which is not a criticism of them: observer descriptions genuinely do predict job performance across thousands of people and that finding is robust. What does not follow is that a single score licenses a conclusion about the single person it came from, because the individual in front of you may sit anywhere in the distribution, including at the end where the relationship reverses.
Almost the whole applied industry runs on treating the second as though it followed from the first, and so, as the next part shows, the automated version does too. The difference between a defensible use of any of this and an indefensible one is not the quality of the instrument, and it is not whether the observer is human. It is whether the decision being taken is the kind the instrument can support.
PART III. THE OBSERVERS NOBODY AGREED TO
Chapter 6: The Machine Is Also in the Room
6.1 Asked For, or Taken
Everything so far concerns observation as a social act, built from acquaintance, memory and conversation. Systems now pursue the same object by a different route: converting traces into traits. Two kinds are worth separating before discussing any evidence.
Active observation asks you for something. A questionnaire, a recorded interview, a voice sample, a game. You know you are being assessed, even if you don't know how.
Passive observation collects what you produce anyway. Location, keystroke rhythm, application use, posting patterns, heart rate, sleep, response latency, calendar density. The second is the unsettling one, because you may never experience yourself being assessed while a description of you is assembled continuously.
6.2 What a Phone Can and Cannot Infer
The pooled evidence here is more modest than the coverage suggests. Combining 21 studies, associations between smartphone data and the five broad personality traits were small to moderate, with extraversion clearly the strongest and the other four lower and closely bunched (Marengo, Elhai & Montag, 2023).
Three details from that paper matter more than the headline. Combining many features improved prediction for four of the five traits, so the useful signal comes from patterns rather than from any single revealing behaviour. Call and text records were particularly informative for extraversion, which is unsurprising, since the trait is largely about wanting contact. The authors also found evidence of publication bias, with smaller studies reporting larger associations.
Their conclusion is honest, and it is not the one that gets quoted. This kind of prediction works at the group level, and its use for accurate assessment of an individual is currently limited. That is the same population-and-person distinction as 5.4, arriving from the other direction.
6.3 The Study That Beat Your Friends
The comparison that changed the conversation is now eleven years old and remains unchallenged. Across 86,220 volunteers who completed a hundred-item personality questionnaire, a model working from Facebook likes alone matched self-reported personality more closely than judgements made by the participants’ own Facebook friends, showed higher agreement between judges, and predicted several life outcomes including substance use, political attitudes and physical health better than the human judgements did (Youyou, Kosinski & Stillwell, 2015).
Read that carefully, because two things are true at once. The finding is real and old, which means every discussion proceeding as though machine inference were speculative is a decade out of date. And the criterion is the same one that limits human research, since a model that agrees with your questionnaire has agreed with your questionnaire.
Language has since made this cheaper. Asked to infer the five broad traits from Facebook status updates with no training for the task, large language models reached a level similar to systems built specifically for the job, and did so less accurately for men and for older users than for women and younger ones (Peters & Matz, 2024). The unevenness is the key takeaway. A method can be accurate on average and worse for you.
6.4 Reading the Three Traits From Your Posts
The work has now reached this series’ own subject. A 2026 study compared seven machine learning methods for predicting narcissism, Machiavellianism, and psychopathy from the language of Facebook status updates, and found that one method outperformed most of the others, while all showed small and similar prediction error (Leberecht, Nedderhoff, Zitzmann & Hecht, 2026).
Now notice what that paper actually reports, which is prediction error rather than accuracy. It is a methods comparison, asking which technique fits best, and it does not claim that any of them identifies anybody. The popular coverage announced that artificial intelligence can tell whether you are a psychopath or a narcissist from your social media posts.
That gap between what was measured and what was announced is the single most useful thing in this chapter, and it is not a failure of the researchers. It is what happens when a population-level model meets a headline, and it is the same conversion that turns a questionnaire score into a page of prose about you.
Chapter 7: The Observer That Never Goes Home
7.1 What Is Being Collected at Work
The workplace version is not futuristic. Monitoring can include tracking calls, messages and keystrokes, taking screenshots, recording webcam footage and audio, and using dedicated software to log activity, and around 60 per cent of large employers now use tools that track their workers (Information Commissioner’s Office, 2023). Public feeling about all of this is not remotely ambiguous. Seventy per cent of people surveyed said they would find being monitored by an employer intrusive; monitoring of personal devices was the practice considered most intrusive, and fewer than one in five said they would be comfortable taking a job knowing their employer would monitor them.
The apparent precision of these measures conceals how little they establish. Mouse movement is not productivity, message sentiment is not commitment, and calendar density is not leadership. A system attending continuously to Mary’s response times has no way of knowing about the insomnia, the caring responsibility, or the manager who copies four people into every request, which is exactly the situational information Chapter 5 says the interpretation requires.
7.2 The Interview That Scores You
Automated video interviews are the clearest case, because they are an observer report built by machine for a decision. Researchers built models assessing the five broad traits from verbal, paraverbal and non-verbal behaviour in mock video interviews across several samples, then tested them on a further sample. Models trained on interviewer ratings showed mixed evidence of reliability, consistent relationships with related and unrelated measures, and predictive validity for academic outcomes. Models trained on candidates’ self-reports showed little evidence of either reliability or validity (Hickman, Bosch, Ng, Saef, Tay & Woo, 2022).
That is a more interesting result than either the vendors or the critics tend to report. It says the approach is not empty, that what the model is trained to predict determines whether it works at all, and that the version trained on the cheapest available target performs worst. A buyer who cannot get an answer to what a system was trained on has not been told what decides whether it works.
7.3 Where the Law Has Already Drawn a Line
Two regulatory facts belong in any honest treatment of this, because they show where societies have decided that capability is not permission. Since 2 February 2025, the European Union has prohibited placing on the market, putting into service, or using an artificial-intelligence system to infer a person's emotions in the workplace or in education, except for medical or safety reasons. Emotion inference is one of a small set of practices banned outright rather than regulated, and systems used in employment and worker management are otherwise treated as high risk (Regulation (EU) 2024/1689).
In the United Kingdom, the route was data protection law. In February 2024 the Information Commissioner’s Office ordered Serco Leisure, Serco Jersey and seven associated leisure trusts to stop using facial recognition and fingerprint scanning to record the attendance of more than two thousand employees across thirty-eight sites, and to destroy the biometric data, on the ground that identity cards or fobs would have achieved the same purpose. The regulator’s reasoning included the point that employees had not been offered a workable alternative and, given the imbalance of power, could not meaningfully refuse. The Information Commissioner’s line on why biometric data is different is the one worth keeping: you cannot reset a face or a fingerprint the way you can reset a password (Information Commissioner’s Office, 2024).
The principle underneath both is what this article has been building toward. A system’s ability to make an inference is not permission to make it. The question shifts from whether a description contains signal to whether the observer is entitled to collect it and the decision-maker is entitled to use it here.
7.4 What You Are Entitled to See
This section is for the reader who is not deciding anything about anybody and is simply being described. In the United Kingdom, you can make a subject access request, and it covers personal information collected through monitoring as well as obvious files. An employer that monitors is expected to have a defined purpose, a lawful basis, the least intrusive means available, and, where the monitoring is likely to be high risk, a completed impact assessment; keystroke monitoring and the use of biometric data are given as examples of high-risk processing (Information Commissioner’s Office, 2023).
I would not oversell what that produces. It is a right of access rather than a right of explanation, so it will not always tell you how an inference about you was reached or which signals produced it. It will usually tell you that one exists, which most people do not know, and that is the difference between being able to answer and not knowing there is anything to answer.
The equivalent move outside work is duller and more effective. Look at what a platform’s advertising settings say it has inferred about you, since the inferred-interest list is the closest thing most people can see to a machine-written description of themselves, and it is frequently both wrong and revealing.
PART IV. BEFORE ANYBODY ACTS ON THIS
Chapter 8: The People You Will Wrongly Convict
The commonest error here is not missing a warning. It is acting on something that was never evidence. These are the cases where that happens, written as people because that is how they arrive.
8.1 The One Everybody Finds Difficult for Reasons That Are Not Conduct
Somebody whom colleagues describe consistently and unfavourably presents a pattern that looks like corroboration, and that consistency does persuasive work it has not earned. Descriptions converge because everybody is watching the same behaviour, and they also converge because everybody applies the same expectation, which happens most when a person stands out from the group in a way unrelated to how they behave. The review integrating this literature treats social categories and stereotypes as mechanisms any account of shared perception has to include, alongside accuracy and impression management (Connelly & McAbee, 2024).
The test that separates the two is whether the accounts contain incidents. Genuine convergence produces different people describing different occasions, usually with details that do not match perfectly. The other kind produces the same adjective from everybody with no story attached.
8.2 The New Person Nobody Has Had Time to Read
Six weeks in, the room has a view, and on this evidence the room should not have one yet. A judgement formed from six weeks of frequent contact is made in exactly the condition Chapter 3 identifies as weak, and early impressions become the frame that later evidence is read into.
8.3 The One Whose Accuser Wants Something
An account of a colleague comes from somebody with a position and often an interest. The research on talk about absent people finds it does several jobs at once, of which passing on accurate information is only one (Wax, Rodriguez & Asencio, 2022).
That does not make the account false, which is the trap in the other direction. A person can be aggrieved and correct at the same time. The interest tells you to check rather than to dismiss, and what you check is whether anybody without that interest describes the same occasion.
8.4 The One Whose Score Sits on the Other Slope
This one is specific to trait reports and it is invisible without 5.2. Two people with the same score are not the same case when the relationship between the score and the outcome rises and then falls, and a report showing a position on a scale does not mark the turn.
Anybody deciding from a numerical profile is therefore making it from a figure that is ambiguous in a way the document does not disclose. The correction is not a better report. It is to stop using the number as the basis of a decision about the individual.
8.5 The One the Model Got Wrong Because of Who They Are
Automated inference is not evenly accurate. Language-model inference of personality from posts was less accurate for men and older users, and that unevenness is a general property of these systems rather than a defect of one of them (Peters & Matz, 2024). As a result, a person can be misdescribed by a system precisely because they belong to a group the system reads less well, and nothing in the output will indicate that this happened. A confidence figure, when provided, describes the model rather than the case.
8.6 The One Whose Context Collapsed
A remark made for one audience travels to another without the knowledge that made it intelligible. A model reading posting patterns cannot distinguish a bad year from a bad character, and neither can a colleague reading calendar density, which is the same failure in two forms.
8.7 The One Who Needs a Clinician, and the Person the Numbers Do Not Describe
Nothing here supports labelling anybody with a condition. Several people agreeing about a colleague is not a clinical judgement and cannot be converted into one, and if somebody’s conduct suggests they are unwell, that is a matter for occupational health and a qualified clinician with direct access to the person.
And every figure in this article describes tendencies across groups. Not one of them describes the individual you have in mind. That caution is worth restating here because this article is most likely to be used as ammunition.
Chapter 9: A Fairer Way to Be Known
Each item carries its evidence status. Evidenced means supported by the research above. Regulatory rule means it comes from a regulator rather than a study. Practitioner judgement means it comes from my own executive and advisory practice and is labelled as such, because the research on what to do is far thinner than the research on how any of this works.
9.1 If You Are Judging Somebody
Ask people who know them rather than people who have seen them, because closeness, not exposure, drives accuracy, and the two are usually different people. Evidenced. Establish how many genuinely independent accounts you have before deciding what any of them is worth, since a substantial share of any single account belongs to the person giving it, and several accounts that share an origin sit closer to one than to several. Evidenced.
Ask what somebody saw rather than what they think, every time, because a verdict arrives already blended with the judge, while an incident can be dated, checked and answered. Practitioner judgement, on an evidenced foundation.
Ask what the situation was paying for. If the behaviour you are judging was rewarded by the role, you are looking at an environment as much as a person, and moving the person without changing the reward reproduces the problem. Practitioner judgement.
Put the same four questions to an automated score that you would put to a referee: what was it trained to predict, against whose ratings, how stable is it, and for whom does it work less well. A supplier who cannot answer the first two has not given you an instrument; they have given you an output. Practitioner judgement, on an evidenced foundation.
And do not run a consequential judgement about a person through a single manager, or a single model, which is the same defect wearing different clothes. Evidenced.
9.2 If You Are the One Being Described
Start by finding out what exists at all. In the United Kingdom, a subject access request reaches personal information collected through monitoring, and an employer carrying out monitoring likely to be high risk is expected to have completed an impact assessment. Regulatory rule.
Know where the lines already are. Inferring a worker’s emotions through such a system is prohibited in the European Union outside medical and safety uses, and a UK employer was ordered in 2024 to stop scanning employees’ faces for attendance and to destroy the data, because a card would have done the same job. Regulatory rule.
Separate what you did from what you are and answer the first. A dated incident can be explained, contested or accepted. A label cannot be any of those things, which is why the moment one is applied, the conversation stops being about conduct that could change. Practitioner judgement.
Weigh the sources you are given by the same rules you would apply to somebody else. A pattern named by one person who has known you for fifteen years is a different object from the same adjective repeated by four colleagues who share a group chat. Evidenced.
And treat your own account as evidence rather than as the truth. The one finding in this article that lands directly on the reader is that people who overestimated their agreeableness and conscientiousness performed worse than those who did not. Evidenced.
9.3 Three Refusals
Do not reach for a questionnaire to settle a suspicion, because asking somebody about themselves is the weaker of the two available routes on these traits and the published measures are research tools distributed for non-commercial use. Do not put a label on a named person, no matter how many accounts agree. The output of everything described here is a better-sourced description of behaviour, and never a category for a human being.
Do not build a back channel. A network that collects accounts without asking what was seen produces reputations rather than evidence; it is most easily steered by the person best at steering it, and its subjects cannot answer what they are never shown.
Chapter 10: What This Changes
10.1 The Short Version
Other people’s descriptions of you carry real information, including about the traits you would least like seen. Pooling 24 studies and 7,022 people, self-descriptions and descriptions by people who knew them converged at a level above what ordinary personality traits usually produce.
Closeness, not exposure, makes an observer accurate. Across 263 samples and 44,178 people, intimacy was needed for the substantial gains, making family and close friends better witnesses than colleagues and making the workplace a setting of much observation and modest accuracy.
The science of how people are known is domestic, and employers borrowed it. Two of those twenty-four studies used people at work, and the borrowing dropped everything that made the findings interpretable, leaving a questionnaire and a page of prose.
A score cannot tell you about a person. Trait relationships are not all straight lines, the same number can sit on either side of a turn, and behaviour depends on whether the situation rewards it.
A model reading your likes matched your self-description better than your friends did, eleven years ago. Machine inference is neither speculative nor impressive: it works at group level, unevenly across groups, and with no access at all to the circumstances that produced what it is reading.
10.2 So What Do You Do
If you decide about people, move the weight. Towards descriptions by people who know them, towards incidents rather than adjectives, towards several independent sources, and away from any single number, whether a manager produced it or a model did. If you are being described, find out what exists, focus on conduct rather than character, and use the access you already have. Most people do not know that the monitoring data is reachable, and that single fact is the most immediately useful thing in this article.
10.3 Why Any of This Matters
The organisational cost of getting this wrong is not the bad hire. It is the two years in which everybody around a difficult person knew, said so to each other, and correctly judged that saying so upwards would achieve nothing. The personal cost is different and larger. It is that the number of entities forming a view of you is rising, the proportion of them you can question is falling, and the descriptions that will decide things are increasingly produced by observers with more data about your behaviour than anyone has ever had and no knowledge whatever of your life.
10.4 What I Still Do Not Know
The largest gap is the one at 2.4. Almost every result here compares a description against the person’s own account, and where those agree, the research records a success without establishing that either is correct. Second, machine literature moves faster than it validates. The strongest comparison with human judges is from 2015: the smartphone pooling reports publication bias, and the dark-trait modelling work reports prediction error rather than any claim about identifying individuals.
The third is that I could find no research on the event this article opens with: a pattern named informally by somebody who knows you well and discounted by somebody who does not. It is the commonest way accurate information about character moves between people, and it appears to have no literature.
And I do not know what happens to any of this when the descriptions become continuous. Every finding here concerns an episode of judgement. Nobody has yet studied what it does to a person to be described without interruption by something that never forms an opinion, only an output.


