Love

Can the Spark Be Calculated?

I tested harmony by algebra.

— Alexander Pushkin, “Mozart and Salieri”

The figure “94 percent” sounds more convincing than “you may well have something to talk about.” Samantha Joel, Paul Eastwick and Eli Finkel stripped that promise of its convenient vagueness. They collected speed-dating data — more than a hundred characteristics per person — and set the machine a concrete task: to predict, before the meeting, how much these two particular people would be interested in each other after talking.

The machine learned two things. It could predict general attractiveness: some people were liked by many. And it could predict a general tendency to be taken with others: some participants gave high ratings to almost everyone. The third task proved more stubborn: whether a particular A would like a particular B more than their general tendencies would lead you to expect. That is what a user usually calls compatibility. On the unique pairwise part of the rating, the model barely moved.

Testing on people the model had not seen

The main test began with people whose answers the machine had not seen during training. On a familiar sample, a complex algorithm can memorise chance combinations of music, humour and ratings and take noise for a pattern. So part of the observations was held back until the end. On new participants, the model confidently distinguished general attractiveness and a general tendency to be taken with others, but it barely predicted the unique pairwise part — the particular pull of A towards B.

After a good date, an explanation is easy to find: they both love dogs, agree about politics and laugh at dry jokes. Such an account can be accurate and still fail to separate this couple in advance from dozens of others with the same overlaps. A cause that convinces after the event does not thereby become a good predictor before it.

The line “the algorithm knows you better than you know yourself” needs the object of that knowledge spelled out. Your click history — possibly. The chance you will answer a message — sometimes. The unique pairwise reaction before any interaction is predicted considerably worse by the available research. This is not the limit of every possible algorithm: video, correspondence and observed interaction supply data of a different order. The more precisely the goal is stated, the less magic is left in a compatibility percentage.

Three components of a rating

Suppose that after a meeting a person gives his partner seven out of ten. That number can be broken apart. The first part belongs to the partner: how highly other people usually rate him. This is the general component. The second belongs to the rater: how generous his scores are overall. The third belongs to the pair: what happened between these two beyond their average tendencies. The first two components are relatively easy to compute. The third requires information about the interaction. Questionnaires contain data about people separately and almost no data about a couple that did not yet exist.

The algorithm was shown two photographs of ingredients and asked to name the taste of a dish that had not been cooked. The study leaves other sources of prediction open: video of a conversation, features of speech, synchrony of movement or early correspondence may add useful information. But detailed questionnaires from two people before they meet reveal almost nothing on their own about their mutual reaction.

The data appears after the meeting

For an algorithm, the more important consequence of splitting the general rating from private taste is that the object of prediction changes as people get to know each other. Before the conversation, what is available is a face, clothes, status and a questionnaire; after several meetings, pairwise data appears — shared jokes, ways of repairing awkwardness, a sense of safety and the predictability of a reaction, none of which existed when the profile was filled in.

For an industry built on precise matching, the awkward part is this shift in the level of the data. In an analysis of 43 longitudinal studies, current relationship quality was best associated not with general personality traits but with what had already arisen inside the couple: the partner’s perceived commitment to the relationship, appreciation, sexual satisfaction, conflict and the sense that the partner was involved. Changes in quality over time were predicted far worse. An algorithm starts to see a union clearly after the union has produced a history of its own; before the meeting, that history does not exist.

A prospective study of early relationship development followed 208 participants and 1,065 potential partners over seven months. Romantic interest was better explained by judgements formed about a specific person: a sense of compatibility, perceived attractiveness and other pairwise impressions. Matching a previously stated ideal, and combinations of stable individual traits, added substantially less. A model begins to see the pairwise reaction after the first interactions appear, not after the questionnaire gets longer.

The algorithm as a transport system

Apps are strong as a transport system. They filter out incompatible goals and distance, take a person beyond his usual circle, and use behaviour on the platform to predict who will reply or keep the conversation going. That data is closer to clicks and messages than to a couple’s future life. The user is looking for closeness, trust and a durable relationship; the platform observes activity. Optimising the second can keep a user on the platform with the help of people who are more interesting to scroll through than to live with.

A compatibility percentage often does practical work regardless of its accuracy: it gives two strangers a reason to start talking. That does not require being an oracle.

After the first contact, the search should change instruments quickly. Correspondence tests politeness, goals and the ability to keep a conversation going. A meeting adds a bodily reaction and basic reciprocity. Repeated meetings show whether behaviour is consistent. Each stage produces data of a different quality, so a profile only answers the question of whether there is enough reason to have coffee, not the question of whether to build a union.

The figure of 94 percent can be a good invitation to coffee. The percentage becomes misleading when it is read as the result of a coffee nobody has had yet.

Main sources

Joel, S., Eastwick, P. W., & Finkel, E. J. (2017). Is romantic desire predictable? Machine learning applied to initial romantic attraction. Psychological Science, 28(10), 1478–1489.

Eastwick, P. W., & Hunt, L. L. (2014). Relational mate value: Consensus and uniqueness in romantic evaluations. Journal of Personality and Social Psychology, 106(5), 728–751.

Joel, S., Eastwick, P. W., Allison, C. J., et al. (2020). Machine learning uncovers the most robust self-report predictors of relationship quality across 43 longitudinal couples studies. Proceedings of the National Academy of Sciences, 117(32), 19061–19071.

Eastwick, P. W., Joel, S., Carswell, K. L., Molden, D. C., Finkel, E. J., & Blozis, S. A. (2023). Predicting romantic interest during early relationship development: A preregistered investigation using machine learning. European Journal of Personality, 37(3), 276–312.