A Field Under Many Names

Five literatures that barely cite one another are converging on the same object. A category is forming, and it is forming now because the data substrate is new and because one representation of a person makes downstream tasks possible that did not previously exist.

Platform behavior Life and health Transactions Cognition Interface logs one representation of a person Recommend Assist Simulate Match Forecast Experiment
Fig 1 · The shape of the category. Five literatures pretrain on different records of behavior and induce the same object. What makes it a foundation model rather than a technique is the right-hand side: many tasks read from one representation, and two of them had no commercial form until recently.

For the past several months we have been reading outside our lane, and the same thing kept happening. A paper from recommender systems, a paper from computational psychology, and a paper from clinical informatics would describe what looked like one research program in three vocabularies that had no words in common. None of the three cited the others. Each seemed to believe it was solving a local problem.

We have collected that reading into a public list, awesome-foundation-models-of-human-behavior, covering models pretrained on records of what people do, whose learned representations of a person transfer across tasks. It is open and we would welcome corrections and additions.

Assembling it changed what we thought we were looking at. This is not several fields that happen to rhyme. It is one category coming into existence, and it is arriving now for two specific reasons: a substrate of behavioral data that did not previously exist at this scale, and a set of downstream tasks that only become possible once a single representation of a person is good enough to support them.

Five literatures, one object

The list is organized by where the behavioral data comes from, which turns out to be the axis along which the field actually split.

LiteratureRepresentative work
Platform behavior Meta's HSTU recasts ranking and retrieval as generative sequence transduction over user action streams. Snap's UUM does multi-task next-event prediction serving both ads and content. OpenOneRec is the first major open-weights release of the genre.
Life trajectories and health life2vec treats Danish national registry records as a language of life. Delphi-2M forecasts more than a thousand diseases roughly twenty years out. Both appeared in Nature.
Money and transactions Visa's TransactionGPT and Sber's CoLES and LATTE learn representations from payment streams, with LATTE aligning transaction embeddings to language-model descriptions of behavior.
Cognition and simulation Centaur is fine-tuned on ten million choices from psychology experiments. Socrates draws on 2.9 million responses across 210 studies. OdysSim trains on 21.4 million interactions across 62 datasets.
General user models GUM builds user models from screenshots with confidence-weighted propositions. LongNAP predicts next actions from screen logs. PTUM stated the "pretrain a user model like a language model" idea back in 2020.

An economist studying transaction streams, an epidemiologist studying registry data, and an ads engineer studying click sequences would not normally read each other. But the artifact each produces has the same four properties. It is trained on behavioral traces rather than on descriptions of people. It is trained under self-supervised pressure, predicting the next thing a person does, which requires no labels and so scales to all recorded behavior rather than the sliver someone annotated. It induces a representation of an individual person. And that representation is consumed by more than one downstream task.

Those four clauses are checkable. Applied to the papers above, they pick out the same set regardless of which field the paper came from, which is the strongest evidence we have that this is one category rather than a metaphor connecting several. The convergence is not an argument anyone had to win. It is visible on inspection once you line the systems up.

Why now: the data is a different substrate

Categories in machine learning tend to open when a substrate becomes available, not when someone has a good idea. Language models became possible when the web made text abundant. What is opening this category is that records of behavior have become abundant in the same way, and behavior is a genuinely different kind of data from text about behavior.

The distinction that matters is between the stated self, the revealed self, and the valued self. Almost everything the industry currently builds on is the stated self: profiles, surveys, preferences typed into a settings page, a paragraph in a system prompt. It is cheap to collect and it is the weakest of the three, because what people say about themselves is a document they authored, not evidence of what they do. The revealed self is the trace: what was actually clicked, bought, opened, abandoned, returned to. The valued self is what someone would endorse on reflection, which is often neither of the other two, and which almost nobody has figured out how to instrument at all.

Behavioral traces also carry something text does not. Because they are logged rather than composed, they contain the choices people would not narrate, including the ones they would not admit to. That is what makes them powerful and it is also why the collection question cannot be treated as an afterthought. The models in the list above were mostly trained on data that already existed for other reasons, ad logs, medical registries, payment rails, which is why every one of them stops at an institutional boundary.

So the open question in this category is not architecture. It is what a purpose-built behavioral dataset looks like, gathered with consent, spanning more than one silo, and structured so that a representation learned on it survives leaving the context it was collected in. Nobody has that yet. It is the part we think is worth building.

What one representation makes possible

The other half of what makes this a category rather than a technique is the task side. A foundation model is only a foundation if many things can be read off it, and here the list of things is longer than it looks.

Six tasks come off the same representation. Recommend, which the platforms already do well. Assist, meaning an agent that acts on your behalf without being briefed each time. Simulate, running a synthetic population to see how people respond before you ship. Match, connecting people to people or to opportunities. Forecast, projecting what someone will do over a long horizon, which is what the health trajectory work is already demonstrating. And experiment, asking a counterfactual question about a population you cannot ethically or affordably run a trial on.

Several of these are not currently products. Simulation and experimentation in particular have no mature commercial form, because until recently there was nothing faithful enough to simulate against. They become available not because someone built a simulator but because a representation good enough to support one now exists. That is the characteristic signature of a foundation model category: the interesting applications are the ones nobody could attempt before, and they arrive together because they share a substrate.

It also means the category should be measured differently. A model that is excellent at one of these six and useless at the other five has not demonstrated generality, it has demonstrated a good recommender. Generality is a property of the profile across tasks, and a field without a shared way to measure that profile cannot tell the two apart. That is the piece of infrastructure most conspicuously missing right now, along with a behavioral vocabulary that survives crossing from medical events to purchases, and a legal path for data to pool at all.

The strongest objection

We should state the case against, because it is a real one and it is not yet settled.

It is possible that what looks like a model generalizing about a person is closer to token-level memorization of that person's recorded history, dressed up in the vocabulary of representation. Relatedly, nobody has yet demonstrated that training on behavioral data improves a language model's understanding of people in a way that survives contact with a new population. Both objections are testable on public artifacts, and until someone runs those tests the field is holding a promissory note.

There is a measurement problem underneath this. To say a representation falls short, you need to know the ceiling, and for most tasks about people the ceiling has never been measured. Compatibility prediction is the sharp example. Out-of-the-box models are plainly insufficient at predicting who will get along with whom, but no test-retest ceiling for that task has ever been established, so nobody can say how much of the shortfall is fixable and how much is intrinsic to people being inconsistent. Conflating a hard task with a bad model is the most common evaluation error in this space, and we have made it ourselves.

The category is being priced

While the literature stayed fragmented, capital started naming the thing. In July, Simile raised more than $200 million at a $2 billion valuation, led by Greenoaks with Index increasing its position, five months after a $100 million Series A. The company describes itself as building a foundation model that predicts human behavior at scale, and its founding team comes out of the Stanford group behind the generative agents and social simulacra work that appears on our list.

What makes it a useful datapoint is not the number. It is that the company is being valued on simulation, one of the tasks that did not have a commercial form until very recently. A market forming around the newest of the downstream tasks, rather than around a better recommender, is what category emergence actually looks like from the outside.

What we are doing with it

We are writing a survey of this space. It works through the definition above, tests it against the largest literatures, takes the objections seriously enough to specify experiments that would settle them, and ends with an agenda: a standing benchmark that measures across tasks rather than within one, a behavioral vocabulary that survives crossing a domain boundary, a consented data path, and predictions with dates attached so the field can be wrong in public.

It is a working draft and not ready to share yet. The reading list is, and it is the part most useful to other people right now. If you work on any of these five literatures and think we have mischaracterized your corner of it, or if we have missed something, the repository takes issues and pull requests.