Scoremaxxing: The feature that didn’t exist until users showed up
B2C apps have to give people a real reason to download them and actually keep using them. Cracking the unique value- the thing that fits your specific audience- is the whole game. Get that right and you have an app. Get it wrong and you have a download number that decays week over week.
And scores. Oh, scores.
We see them everywhere. Your Spotify age. Your completion score for an amp workout. Your sleep score in Apple Health. They’re more than just a number- they’re calculated data that reflects something specific about you in the context of that one app. Your music taste, your training, your recovery. Each score is a tiny argument the app is making: here is something true about you that you didn’t know before you opened me.
And then there’s community. The amp Fitness Score plays in that space too.
The whole point is to tell you where you sit in a distribution of other users- and then adjust your next workout based on that. Working weights, rep targets, progression curve, all computed from where you land in the population.
Which is a problem, because on day one there is no population.
You can’t test a percentile engine without a distribution. You can’t validate age-adjusted curves without users of different ages. You can’t tune the prescription logic without watching a real human get a real recommendation and either crush it or struggle. So you do what every engineer building a data-dependent feature ends up doing: you fake it.
We built a synthetic population. Fake users at different ages, training histories, body compositions, plausible force curves. Enough scaffolding to exercise the math and make sure we weren’t about to recommend something insane to anyone.
But mocked data has a ceiling. It tells you the code works. It can’t tell you the feature works. Because the feature is the data.
Then we shipped, and something quietly cool started happening. Every completed benchmark sharpened the distribution a little. The curves got tighter. The prescriptions a new user got were calibrated against a dataset that didn’t exist 2 weeks ago before calculation.
And here’s where it gets fun- because the benchmark isn’t the whole story.
The benchmark score is one of three inputs to something bigger: the amp Fitness Score.
Benchmark: where you sit in the population on a standardized protocol.
Consistency: are you actually showing up.
Effort: when you do show up, how hard are you working.
The Fitness Score is the composite. One number trying to honestly answer the thing every user actually wants to know: am I getting better?
This is the part I’m most into. Two people with identical benchmark scores can have very different Fitness Scores, because one trained four times a week and went hard, and the other showed up twice and coasted when his goal was 3 workouts a week. The score is a fingerprint, not a leaderboard.
And it’s not static. It gets more meaningful the longer you use the app, because it has more of you to work with. Day one it’s a guess. Month three it’s a portrait.
The thing I keep coming back to is that none of this works without trust. Personalization at this level only happens if users hand over real data about their bodies, their effort, their consistency. That’s not a small ask. The responsibility on our side is to use it wisely- keep it private, keep it secure, and make sure every calculation we run on it actually gives something back to the person it came from. No selling it, no leaking it, no using it for anything other than making their experience more theirs.
That’s the deal, and honestly it’s the part of the work I’m most into. Data used carefully, with respect to user’s privacy, but USED, is what makes an app feel like it knows you.

