When AI isn’t so I: Health Data Reliability and Validity

By Karen Sternheimer

author photo

As a social scientist, I like data. If you can quantify something, I’m interested. Data is the foundation of sociology: we are not making casual observations and calling it science; we empirically gather data and look for patterns.

You might have learned about the concepts of validity and reliability in a research methods class. Validity requires that something actually measures what we say it measures. For instance, it’s clear that stepping on a scale is not a valid measure intelligence. Reliability requires that our measure is consistent. Using the scale analogy, it should give us the same result if we step on it a few minutes later. If not, the scale is off.

As someone very interested in health, I have been wearing a smartwatch primarily for fitness feedback for the past 6 years. I have learned more about the limitations of Artificial Intelligence (AI) in this process than anticipated. It’s a good reminder when we get caught up in a panic over the future of AI that it might be artificial, but it isn’t always very intelligent. I won’t say which watch I wear but suffice to say it is a high-end version of a very well-known brand.

Let’s start with my sleep score. No, I don’t need a device to tell me how I feel after waking up in the morning and whether I had a good night’s sleep, but I’m often curious to see what the watch says. It purportedly uses heart rate and motion data to determine when I’m asleep, and what sleep cycle I’m in. It’s not clear that the proprietary algorithm is producing valid measures of any of these things since I haven’t done a formal sleep study with electrodes to monitor my brain activity while asleep.

My watch gives me a score out of 100 each night. Sleep duration counts for 50 points, bedtime 30, and interruptions 20. I regularly get a low interruption score even when I have slept soundly because I turn over a lot.

I am extremely consistent with my bedtime, but I will get a deduction for going to bed 10 minutes later than my average time on, even on a Saturday night. A human would probably be shocked by how early I go to sleep—especially on a Saturday night—and perhaps give extra points for that level of consistency. A human might also figure that 10 minutes really doesn’t matter, but I will get chided by the app and lose points for missing my exact bedtime. I know this isn’t a valid measure of how well rested I am, but I look at it every day anyway.

Heart rate is a very important measure for training; the fitter one gets, the lower your heart rate should be for the same tasks. For years, I didn’t realize how my watch was not reliably measuring my heart rate. It was telling me an easy run was an “all out” effort at my maximum heart rate, claiming I was staying in this top zone for hours at a time, which is just not physically possible.

After searching online for a solution, it turns out that when running I need to wear the watch higher up on my arm to get a more reasonable (and hopefully more accurate) heart rate. When doing other activities, wearing the watch closer to my wrist seems to work just fine. Besides a problem of reliability, the watch was apparently confusing my running cadence, which it measures while running but not during other activities, with my heart rate.

The watch’s “brain” is too limited here to know the difference. This interferes with other measures which may have little validity, such as my VO2 max, or total aerobic capacity, and my cardio recovery, which plummeted since I fixed the watch to properly read my heart rate. But these measures are just the result of AI estimates; to get valid results they need to be measured in a lab.

The app constantly wants what it measures to trend up, even when it makes no sense for it to do so. A few months ago I trained for and ran a marathon, and since then my average daily mileage has decreased as I have recovered. The app has a message, “Your total mileage is down, Karen. If you want to get this arrow back on track, try to find time for a walk each day [emphasis mine].” My current average daily mileage is 8.1, down from 9.1. I’m obviously finding time for a walk and more. I think I’m “on track.”

It has also reprimanded me for my exercise minutes decreasing from an average of 191 to 183 per day, telling me to “Start by averaging at least 192 Exercise minutes” per day this week.

I think it’s fair to say the health app does not give intelligent advice. The American Heart Association recommends 150 minutes of “heart-pumping” exercise per week, and I get more than that in a day. Beyond bad advice, it can be dangerous. It doesn’t consider too much exercise a possible problem. When working with a physical therapist earlier this year, he emphasized the importance of rest days, not trying to arbitrarily maintain mileage or minutes.

AI tools can be useful, but they are not as smart as we are, or at least not always. Spell check and auto correct are good examples of AI tools that can miss context and meaning (my name gets auto corrected to Kate a lot). So are apps that generate essays, which often involve somewhat superficial content scraped from similarly superficial websites.

What AI tools have you noticed in your everyday life? What issues with validity and reliability have you noticed?

Leave a Reply