Research

Research Blog · September 2026

The same trait is a different question in a different room

By Hon-Ming Gianotti and Eric Shamlin

All results below are in-simulation.

Introduction

Suppose you want to know whether someone is an extravert, and all you are permitted to observe is whether they say ‘yes’ to things. It seems a reasonable enough proxy – extraverts are, after all, meant to be the people who say yes. This post is about why that proxy is not merely noisy but, under quite ordinary conditions, points the wrong way, and about what our simulated people showed us before we had thought to look for it. It will first set out the puzzle as it appeared in our own data. Then, it will explain why psychometrics has traditionally treated behaviour of this sort as a nuisance to be detected after the fact. Finally, it will describe the result we find most interesting from the past year of work: that a model in which personality reaches behaviour only through intervening machinery does not merely tolerate this context-dependence, but produces it – and, more usefully, predicts it.

The puzzle

Our research programme at Tiresias builds simulated populations in which each person is an assembly of cognitive mechanisms rather than a single rule, and in which every one of those mechanisms is set by a latent personality profile (the Big Five, plus a general executive-function factor). We will not describe the internals here. It suffices for now to say that a trait never touches behaviour directly; it can only adjust some intermediate process, and it is that process which meets the situation.

When we asked how well each trait ‘discriminates’ on acceptance behaviour – how strongly, that is, a person's standing on the trait predicts whether they say yes – we naturally expected the answer to vary by situation. What we did not expect was the sign to change. For extraversion, discrimination measured on acceptance rates is +0.22 in situations rich in social information and −0.28 in situations poor in it. In socially rich settings, extraverts say yes more; in socially barren ones, they say yes less. The trait did not change; the room did. Openness, meanwhile, discriminates almost exclusively on novel items (+0.35 novel against +0.01 familiar), which is to say that an openness measure built from familiar choices measures nothing at all. And executive function was invisible across the entire surface of acceptance behaviour, for reasons we return to below.

Trait discrimination · Tiresias

The trait didn't change; the room did — discrimination on class-conditional acceptance (r)

Tiresias
−0.40−0.200.00+0.20+0.40ExtraversionOpennessExecutive functionSocial-information-rich+0.22Social-information-poor−0.28Novel items+0.35Familiar items+0.01All acceptance classes0.00
Discriminates toward yesDiscriminates toward noIn-simulation result, research programme v4.6

Why this is normally a nuisance

Psychometricians have a name for an item that behaves differently for different groups of people at the same underlying trait level: differential item functioning, or DIF. In the standard treatment it is a defect. One builds an instrument, administers it, and then runs statistical checks to find items that misbehave across (say) cultures or administration contexts, and one either removes those items or models around them. The detection is always post hoc and (admittedly) somewhat atheoretical – the tests tell you that an item drifts, but rarely why, and never in advance.

What happens when traits must travel

If a trait reaches behaviour only through some intermediate process, then the discrimination of that trait on any behaviour is a property of the route, not of the trait. Extraversion, in our model, adjusts (among other things) how heavily a person weights social evidence. Where there is a great deal of social evidence, that weighting pushes toward yes; where there is very little, the same weighting has – perhaps counter-intuitively – the opposite effect, because a sparse social field now counts against the option. The sign flip is not an artefact we had to explain away. It falls directly out of the route.

Put another way, the architecture is itself an item response function. Treat each class of situation as an item, and the model tells you, for every trait, how discriminating that item is and in which direction, before a single participant has been run. DIF stops being something one detects and becomes something one derives.

The part that surprised us

The result we find most impressive from the past year is what happened when we used this derivation the other way around. For six rounds of experiments we had failed to recover executive function from simulated behaviour cleanly; whatever signal we found kept collapsing into conscientiousness. The DIF surface said why: executive function has no pathway to whether one says yes, so an acceptance-based battery cannot measure it, by construction, however long one runs it. It also said what kind of task would work and – equally usefully – what kind would not.

We built one task to that specification. Recovery of executive function went from 0.317, with a discriminant-validity violation, to 0.537, clean, in a single step. One must note, however, that this is a within-simulation result: the model predicted the behaviour of the model. That is a weaker claim than it would be with human participants, and we do not want to dress it up as more. But it is a prediction that could have failed, and it did not, and it resolved a problem that six rounds of intuition-led item design had not.

Executive-function recovery · Tiresias

Before and after the task the DIF surface specified — held-out recovery r (5-fold CV)

Tiresias

0.3170.537

in a single step — from a discriminant-validity violation to a clean pass

0.00.20.40.6registered pass threshold (0.50)beforeafter
Six prior rounds of intuition-led item design had failed; one task built to the model's specification fixed it in one step.In-simulation result, research programme v4.6

What it would mean if it holds

The hypothesis we most want to test on real data is the obvious extension: that the DIF observed in real instruments across contexts – extraversion items behaving differently in solitary and social administration settings, say, or across cultures with different densities of social information – is predictable from the structure of the mechanisms in between. If so, item design would become derivable from cognitive theory rather than iterated by trial. We would rather state plainly that this is, for now, just a hypothesis. The real-data test is being pre-registered, and we will publish the outcome whichever way it falls, as we have done with every registered test in this programme so far.

In sum

An extravert is not simply a person who says yes more. In our simulated populations, at least, they are a person who says yes more in some rooms and less in others, and a model that knows why can tell you which rooms will measure what before you build the questionnaire. Whether real people behave the same way is the question we are now setting out to answer.