Key points
- โOur internal north-star metric is movement on validated clinical measures at 12 weeks, not daily active users or session length โ DAU is tracked, but it doesn't appear in the metric leadership reviews against.
- โWe use established, validated clinical instruments for outcome tracking rather than a custom in-house score, specifically so our numbers are comparable to outside research instead of only meaningful within our own dashboard.
- โWhen a feature increased engagement without moving outcome measures, we shipped that finding internally as a caution flag, not a win โ engagement and outcome data are reviewed side by side so one can't quietly stand in for the other.
Every product has a north-star metric, and for most consumer apps it's some flavor of engagement โ daily active users, session length, retention curves. Those are the numbers that show up first on any dashboard, because they're easy to measure and they move fast enough to react to. None of them tell you whether anyone using the product is actually doing better, which is a strange thing to leave out of the metric a mental health company organizes itself around.
We track engagement numbers too โ they're operationally useful, and ignoring them would be its own kind of mistake. But they don't sit at the center of how we evaluate whether YouMindo is working. Building a metric that does sit there took longer than building most of the product features it's meant to evaluate, and it's still the metric we spend the most internal argument on, more than a year after we first shipped it.
Choosing a north star that isn't engagement
Our internal north-star metric is movement on validated clinical outcome measures at 12 weeks of use, tracked per client and aggregated across the platform. It's slower to move than an engagement number, harder to attribute to any single feature, and more expensive to collect, since it depends on people actually completing periodic clinical check-ins rather than passive usage data. We chose it anyway because it's the only number on our dashboard that answers the question the company actually exists to answer.
Why we use validated instruments, not a custom score
We deliberately didn't build a proprietary 'wellbeing score' as our primary outcome measure, even though a custom score could be tuned to be more sensitive to exactly what our product does. We use established, validated clinical instruments instead. The trade-off is that a validated measure isn't tailored to us โ but it means our outcome numbers are comparable to outside clinical research and can be checked against instruments clinicians already trust, instead of being a number that's only meaningful inside our own dashboard.
The mistake: a dashboard that only showed engagement
For the first year, the metrics dashboard leadership reviewed weekly was, functionally, an engagement dashboard โ DAU, session length, feature adoption, retention curves โ with outcome data living in a separate, less-visited report generated monthly. That structure quietly taught the whole company to think in engagement terms day to day, because that's what was in front of us constantly. We rebuilt the primary dashboard to show engagement and outcome metrics side by side, in the same view, specifically so neither could be discussed without the other in the room.
The week engagement and outcomes disagreed
A notification change we shipped clearly increased daily engagement โ more opens, longer sessions, higher feature usage across the board. Outcome measures for the same cohort, tracked over the following weeks, showed no meaningful improvement over the control group, and a slightly higher rate of users reporting the app felt like 'one more thing to manage.' We wrote that finding up internally as a caution, not a win, and it directly shaped some of the notification-frequency decisions we've made since.
The self-selection problem
Outcome data is only as good as the people who provide it, and the people who reliably complete periodic clinical check-ins are, almost by definition, more engaged with the platform than the people who don't. That means our outcome numbers likely describe a healthier, more engaged slice of our user base better than they describe people in acute distress who tend to disengage from tracking altogether โ the group whose outcomes we'd most want visibility into. We haven't solved this. We flag it explicitly whenever we report outcome numbers internally or externally.
What we do when the two diverge
When engagement and outcome metrics point in different directions, our internal rule is that outcome data wins the product decision, even when the engagement case is strong. That's a genuinely uncomfortable rule to hold to in practice, because engagement numbers are faster, cleaner, and easier to defend in a roadmap review than a 12-week outcome trend with a smaller, noisier sample. We've held to it anyway on every case we can think of where the two have actually conflicted.
Where attribution gets genuinely hard
The honest complication with any outcome metric is that YouMindo is rarely the only thing happening in someone's life during a 12-week window โ therapy, medication changes, life circumstances, and plenty else all move alongside whatever the platform is doing. We use control-group comparisons and cohort analysis to isolate the platform's likely contribution, but we're careful never to claim more causal certainty than the data supports, especially in anything we publish externally.
External validation, and why it's still incomplete
We've partnered with outside researchers to independently validate parts of our outcome measurement approach, which matters because a company grading its own homework has an obvious incentive problem. That validation work is ongoing and covers only part of the platform so far โ we don't yet have independent validation for every feature area, and we say so plainly rather than implying a level of external scrutiny we haven't actually completed yet.
Publishing the numbers that don't flatter us
We've started publishing select outcome findings internally even when they're unflattering โ features that increased engagement without moving outcomes, cohorts where results were flat or mixed. It would be easier to only surface the wins. We think a metrics culture that only reports good news eventually stops being able to tell the difference between what's actually working and what just looks good on a dashboard, and we'd rather catch that early than find out the hard way.
None of this makes measurement easy, and we're not going to pretend our outcome numbers are as clean as an engagement dashboard's. But engagement was never actually the thing we were trying to build. If someone opens YouMindo every day and nothing in their life gets better, we haven't succeeded just because the dashboard looks healthy โ and building our metrics around that distinction, even when it's slower and messier, is the only version of 'measuring what matters' we're interested in.
Tom Walsh holds a PhD in cognitive-behavioural science from Oxford. He leads YouMindo's research partnerships and outcome measurement programs.