Research
What we have published on representing people, and the literature we are reading to do it.
Writing
- AUG 2026How Much Catalog Fits in a Semantic ID
A semantic ID with three levels of 256 codes has 16.8 million addresses. On a 66,000 item catalog it collides on one item in nine. The gap between those two numbers is the whole design problem, and it is measurable from the collision rate alone.
- AUG 2026A Field Under Many Names
Five literatures that barely cite one another are converging on the same object. A category is forming, and it is forming now because the data substrate is new and because one representation of a person makes downstream tasks possible that did not previously exist.
- MAY 2026The Internet of Agents Needs a Model of You
Project Deal showed AI agents can transact for us. When everyone has 100 agents acting in their name, the load-bearing question is whether any of them actually represents the person who sent them.
- APR 2026Emotion Vectors and the Future of Human Models
What Anthropic's new interpretability research means for foundation models of human behavior
- APR 2026Why General-Purpose Embeddings Fail at Modeling People
And what outcome-trained representations get right
- FEB 2026Local Drift-Adapters
Re-embedding a billion-document corpus to upgrade a model is prohibitively expensive. Eight per-cluster adapters close 70% of the gap to oracle.
- FEB 2026Introducing the Embedding Adapter
Zero-downtime migration between embedding models
- JAN 2026Jean Technologies
Building foundation models of human behavior
- DEC 2025The State of AI Memory
You cannot model a person you have not collected context about, so we built memory first. The model is the meta of memory, the representation that captures not just who someone was, but who they are and who they intend to become.
- OCT 2024General Personal Embeddings
A trusted infrastructure for the age of AI
Papers
- FEB 2026Local Drift-Adapters: Mixture-of-Expert Embedding Translation for Heterogeneous Vector DatabasesPDF
Re-embedding a large corpus to adopt a better model is prohibitively expensive. A single global adapter fakes the upgrade cheaply, but its uniform-drift assumption breaks on real corpora. Eight per-cluster adapters close 70% of the gap to oracle.
- DEC 2025AI Memory: A Comprehensive ReviewPDF
A first-principles account of memory in LLM systems: the Experience, Storage, Recall pipeline, the Memory Frontier tradeoff between retrieval precision and semantic coherence, and a survey of the architectures navigating it.
Open resources
A public reading list of models pretrained on records of what people do, spanning recommendation, health and life trajectories, transactions, cognition, and general user modeling. Issues and pull requests welcome.