DigDeeper · experiments ›all projects

Retention · 2026-09-18

Why they leave after one listen

412 accounts submitted a track, played one or two of the recommendations and never came back. This is what they had in common — the device they were on — and what turned out not to matter at all: the recommendations.

It was not the recommendations. On every measure we can compute — genre agreement, artist spread, tempo distance, whether the rows could even be played, and how close the nearest neighbour sat in the embedding — the people who left got results indistinguishable from the people who stayed. On playability they got better results. What separates them is how much they listened: 1.3 tracks versus 5.3, stopping at rank 2 versus rank 7. And playing one or two tracks is worth no more than playing none at all (16.2% vs 17.2% come back). The one place recommendation quality does bite is a small tail: when the nearest neighbour lands below 0.85 cosine — 3.8% of first searches, mostly non-electronic YouTube links — the return rate falls to 23.7% from 38.3%. And the sharpest split of all is not about music at all: the drop-offs are two-thirds phone users, the people who stayed are a majority desktop, and a phone session runs five minutes against fifteen.
412 drop-offs 926 compared against 4 055 accounts in the retention curves 4 903 with a known device first searches, prod, to 18 Sep 2026

Who is being compared

A pattern among people who left means nothing without people who did not, so everything below is a comparison of two groups, measured on the first search each of them ever ran — the moment that decided it.

Drop-offs — 412

Ran a search, played one or two tracks, all on a single day, and have been silent for at least 14 days. Median visit: 1 minute. 95% searched within ten minutes of signing up.

Stayers — 926

Searched on three or more separate days. Averaging 15 searches and 184 tracks played. Same product, same catalogue, same period.

1. The recommendations they got were fine

This is the finding that mattered most, because it is the hypothesis everyone reaches for first. Measured on the top 10 of each group's very first search:

Measure of the first search's top 10Drop-offsStayersVerdict
Rows sharing a genre with the submitted track62.2%64.9%no difference
Distinct artists among the ten9.279.41no difference
Average tempo distance from the submission5.1 BPM3.5 BPMboth tiny
Rows with no playable preview anywhere4.3%7.4%drop-offs had it better
Nearest-neighbour cosine (mean)0.92950.93981 point apart
Nearest-neighbour cosine (median)0.93600.94311 point apart

Genre agreement counts a recommendation as a hit when it shares at least one canonical genre label with the submission; rows without genre labels on either side are excluded rather than counted as misses.

The playability row is the one worth sitting with. Dead rows — results with no embeddable preview on any platform — were almost twice as common for the people who stayed. Whatever drives people away, a result they cannot hear is not it.

2. What was actually different: how far they listened

Their first searchDrop-offsStayers
Tracks played out of the top 101.295.30
Deepest rank they reached2.07.0
Seconds from search to first play3768
Who played anything from the top 10 at all332 of 412815 of 926

The time-to-first-play row runs the opposite way to intuition: the people who left hit play almost twice as fast. Stayers spend a minute with the list before playing anything — reading it, presumably — and then work their way down to rank seven. Drop-offs play the top result within half a minute, maybe the second one, and go.

Two readings fit this equally well from the data alone. Either the list is not inviting people to keep going past the first row, or the people who stop at row one arrived with less intent in the first place. Telling those apart needs session replays, not SQL.

3. Listening is the thing that predicts coming back

Widening from the two groups to every account old enough to have had the chance to return, the relationship between tracks played in the first 24 hours and ever searching again on a later day is steep and monotone — but only after the fifth track.

0 tracks 17.2% 174 accounts 1–2 16.2% 501 3–5 18.3% 698 6–10 25.3% 704 11–25 34.5% 887 26+ 59.4% 1 091 0% 60% Share who ran another search on a later day · accounts whose first search was 14+ days ago

The first three bars are the same bar. Playing nothing, playing one track and playing five all land between 16% and 18%. The curve only starts climbing at six, and by twenty-six it has tripled. So "played one or two tracks and left" is not a distinct failure with a cause of its own — it is the flat bottom of one continuous relationship, and the threshold that matters sits somewhere around the sixth track.

4. Where recommendation quality does bite: the outlier tail

One measure of the recommendations does track retention, and it is not any of the metadata ones. It is how close the nearest neighbour sat in the embedding — the top result's cosine similarity, which is a direct read on whether the submitted track had anything like it in the catalogue at all.

under 0.85 23.7% 156 accounts 0.85 – 0.90 29.9% 499 0.90 – 0.93 30.0% 868 0.93 – 0.96 35.9% 1 619 0.96 and up 38.3% 899 0% 40% Share who searched again on a later day, by the cosine of their first search's top result

Monotone across every step, a 1.6× spread end to end, and it survives controlling for how long each account has had the chance to come back. Notably, genre agreement over the same population does nothing: accounts whose top 10 shared almost no genre with their submission returned at 24.4%, those whose top 10 matched almost perfectly at 28.3%. The metadata is the wrong lens here; the vector is the right one.

What an outlier submission actually is

Comparing the 331 first searches whose top result came back under 0.85 against the 4 954 that came back at 0.93 or better:

The submitted trackOutliers (<0.85)Healthy (≥0.93)
Carries a non-electronic genre (hip hop, pop, rock, latin…)30.5%5.3%
Slower than 100 BPM20.2%2.5%
No genre labels at all28.4%14.8%
Arrived as a pasted URL32.3%28.1%

Six times the rate of non-electronic music, eight times the rate of sub-100-BPM tracks. These are people pasting a music video into an engine whose catalogue is house and techno. The far end of the tail, all of them accounts that never came back:

0.397Simia — Du Faux165 BPM, no genre
0.484Skratch Bastid — Rakim Tribute DJ Set104 BPM, Hip Hop / Funk / Soul, 14 min
0.516Jok'air — Big Daddy Jok (Clip officiel)114 BPM, no genre
0.586Joé Dwèt Filé — 4 Kampé (Clip officiel)92 BPM, no genre
0.632JC NO BEAT, DJ F7, Mc Meno Dani — Maria Mariah130 BPM, no genre
0.650Ina Chansons — Claude François "Alexandrie Alexandra"126 BPM, Disco

Nothing is broken in these searches. The engine answered the question it was asked, honestly, and the honest answer was "there is nothing like this here". The product never says so — it returns a hundred rows ranked 0.4 to 0.6 exactly as it returns a hundred rows ranked 0.96, and the person has no way to tell those apart.

5. What we ruled out

HypothesisWhat the data says
They ran out of creditsDrop-offs spent 1.24 credits and left 7.17 unused. Stayers spent 8.65.
The results would not playDrop-offs got fewer unplayable rows than stayers (4.3% vs 7.4%).
The results were all the same artist9.3 distinct artists in the top 10, against 9.4 for stayers.
The results were the wrong genre62% vs 65% genre agreement, and genre agreement barely moves return rate at all.
The results were the wrong tempo5.1 BPM average distance, against 3.5. Both are within a nudge of the pitch fader.
They submitted weird, long or private files0.7% private uploads, 0.2% over 15 minutes — same as everyone else.

6. One asymmetry worth a second look

How the first search was started separates the two groups more cleanly than anything about the results. Pasting a URL of your own track goes with a 32.7% return rate; picking something out of the catalogue with 22.8%. And on their first search the drop-offs pasted a URL far less often than the stayers — 22.6% against 39.7%.

Not controlled for account age, and the causation could run either way: arriving with a specific track in mind is a sign of intent as much as a cause of it. But it is the largest first-session split in the data, and it is something the product can act on.

7. Phone versus computer — the largest split in the data

Device is the one attribute where the drop-offs stop looking like everyone else. It lives in the analytics rather than the database, so this section is measured against a slightly looser reconstruction of the cohort (229 accounts rather than 412 — the event log's idea of a "search day" is more generous than the database's). Within it the comparison is apples to apples.

Drop-offs 68.5% Everyone else 46.3% Came back 3+ days 42.4% 0% 75% Share on a phone, of the accounts whose device is known

The drop-offs are two-thirds phone users; the people who stayed are a majority desktop. The mobile-to-desktop ratio flips from 0.74 among the returners to 2.2 among the people who left — a bigger separation than anything about the recommendations, the catalogue or the submission.

Holding the observation window fixed and looking only at accounts whose device we know, the whole engagement profile shifts with it:

Signed-in accounts, 14-day windowDesktopPhone
Accounts measured1 1791 072
Median length of the first session15 min5 min
Tracks played in the first 24 hours41.826.8
Never got past two tracks, ever4.4%10.7%
Saved at least one track19.8%10.5%
Dug on from a result18.1%14.4%
Searched again on a later day47.4%41.2%

A phone session is a third of the length of a desktop one, produces a third fewer plays, and is two and a half times as likely to end at the second track — which is exactly the shape of the drop-off cohort. Saving a track, the strongest sign that something landed, happens half as often.

It is not the player

The obvious mechanism would be playback failing on a phone — autoplaying someone else's iframe in a mobile browser is a classic way to lose people. It is not happening. Counting 3.5-second plays against preview starts, mobile converts at 1.07 and desktop at 0.96; no account started a preview and never reached the threshold, on either device, and there are no rage-clicks. Every preview a phone user starts, plays.

So the gap is not a broken feature, it is the shape of the session. Fifteen minutes at a desk with a list of a hundred results is a different act from five minutes on a phone — and the interface is the same one in both cases. Whether a phone-shaped version of digging would hold those people is the obvious thing to find out, and this data cannot answer it.

Coverage: 2 621 of 7 524 signed-in accounts (35%) send only server-side events — an ad-blocker or a declined consent banner means no device is ever recorded for them, and nothing can recover it. They average 99 events against 214 for the rest, so the accounts with no device are systematically less engaged than the ones counted here. Everything above is computed on the 4 903 where the device is known, and the true mobile share of the drop-offs could be higher or lower depending on which way that missing third leans. One row was left out of the table as an instrumentation artefact: phones appear to use filters twelve times as often, because on desktop the filters are always on screen and never fire an "opened" event.

Caveats

Where this points

The lever is not better neighbours for the average search. Those are already as good as the ones that retain. The two things the data actually supports:

  1. Say when the answer is weak. 3.8% of first searches come back with a top result under 0.85 and the interface presents them exactly like a 0.96. Telling someone their track has no near relatives here — and suggesting one that does — is a cheap fix for the one recommendation-quality effect that measurably exists.
  2. Get past the second row. Every account that reaches a sixth track returns at a materially higher rate, and drop-offs stop at rank two after 37 seconds. Whatever makes row three worth reaching is the thing to find, and the place to look is the replays, not the database.
  3. Look at the phone. Two thirds of the drop-offs are on one, their sessions run five minutes against fifteen, and they hit the second-track wall two and a half times as often — while the player itself works perfectly. The same hundred-row list is being served to a five-minute session and a fifteen-minute one.

Built from search_history's stored result snapshots joined against play records, on prod, 18 Sep 2026. The per-account view these numbers were derived from is the Drop-off tool, where each of the 412 can be opened and the recommendations they were given can be played.