What Strava Missed About My Run, and My AI Running Coach Caught

Strava called it my fastest average pace of the month. It said I was building both distance and speed. It found “2 intervals at 7:18/km,” decided the humidity hadn’t slowed me, and described me as energized.

That sounds like a splendid Sunday morning.

I remember something slightly different.

The air felt like “wet, warm cotton.” By the third easy-endurance section, I felt as if I was beginning to lose the plot. By the fourth, I wanted the workout to end. I used conversation and even a little singing to check that the effort was still under control. Some of the slowing and walking was deliberate. Some was because two metal bridges were slick from rain the previous evening. Some was caused by traffic lights.

My chest strap, meanwhile, was producing heart-rate numbers that were physiologically implausible for much of the run.

So which AI running coach analysis was right?

That is the wrong opening question. Strava Athlete Intelligence and my CoachChat process were not given the same evidence or asked to do the same job. Strava produced a useful activity summary from the metrics available inside Strava. CoachChat assessed a prescribed workout using the plan, Garmin evidence, my diary, the weather, RPE, symptoms, breathing checks, sensor-quality review and the chronology of the training block.

The interesting lesson is not that one AI is inherently smarter. It is that context changes what an AI is able to know.

The Run Both Systems Were Looking At

The workout was BT_WK05-03, Easy Endurance Extension, the third run of Week 5 in my BlackToe Holiday 10K build.

CoachChat had prescribed 50 minutes:

5-minute brisk walk

Four 10-minute easy-endurance sections at RPE 3 to 4

5-minute cooldown walk

The goal was not speed. It was to add five minutes to my previous longest session while keeping the running easy and mechanically relaxed.

Garmin recorded about 5.8 kilometres during the 50-minute session, while Strava calculated an average moving pace of 8:22 per kilometre. During the running portions, my Garmin-recorded cadence remained mostly in the high 150s to low 160s steps per minute. I reported pain at 0 out of 10 before, during and after the run. Because of my previous injuries, pain is something I monitor and report especially carefully. My overall RPE was 3 to 4.

Those are useful recorded facts.

The conditions matter too: 20°C, feeling like 23°C, with 93 percent humidity. The route included rolling paved trail, road crossings and slick metal bridges.

That is where a tidy activity card starts to need a witness statement.

What Strava Got Right

Strava correctly surfaced several straightforward facts and comparisons:

The activity covered about 5.8 kilometres.

The displayed average moving pace was 8:22 per kilometre.

It was my fastest average pace of the month within Strava’s comparison window.

Parts of the run were faster than other parts.

The uploaded activity contained an average heart rate around 98 to 99 beats per minute and brief values as high as 161.

That is exactly the kind of quick summary Athlete Intelligence is designed to provide. Strava says the feature uses generative AI to analyze activity data and create personalized summaries. Its launch announcement described the product as translating workout data into simple insights and guidance.

Official Strava information:

Strava’s October 2024 Athlete Intelligence launch

Strava’s February 2025 Athlete Intelligence update

As an activity recap, this was readable, encouraging and easy to scan. I would happily use it as a first look.

The trouble began when the summary moved from describing the recording to explaining the workout.

Where Recorded Facts Became Unsupported Interpretation

Strava called the faster portions “2 intervals at 7:18/km.”

There were no prescribed speed intervals. The workout contained four easy-endurance sections, and I was trying to keep the effort controlled. Pace changes came from effort management, walks, stops, route conditions and the ordinary variation of running outdoors. A faster patch in the file is not automatically a speed interval.

Strava also said the humid conditions didn’t slow me and that I seemed energized.

The app could see weather data and pace. It could not feel the air. I described the conditions as “wet, warm cotton.” The third and fourth work sections became increasingly taxing. Finishing the workout with overall RPE 3 to 4 does not mean every part felt equally easy, and it certainly doesn’t prove that 93 percent humidity had no effect.

The most serious problem was heart rate. Strava classified 91 percent of the activity as recovery-zone work and treated an average around 98 to 99 beats per minute, with brief spikes to 161, as evidence of a controlled easy effort.

The arithmetic may have matched the uploaded values. The physiology did not.

The Heart-Rate Average Was Mathematically Real and Practically Useless

Faulty chest-strap heart-rate trace rejected while running pace and cadence remain stable
A mathematically correct average can still be useless when the underlying signal is broken.

My chest strap failed extensively. The trace jumped from implausibly low values to much higher values, then moved back toward implausibly low readings while pace and cadence remained broadly similar or even faster.

That pattern is not credible as an absolute record of what my heart was doing.

For this run, CoachChat rejected heart rate for:

Zone-compliance claims

Aerobic-efficiency calculations

Cardiac-drift analysis

Any conclusion based on the whole-session average

Any confident interpretation of 91 percent “recovery zone”

This does not prove that my actual effort was too hard. It means the heart-rate file cannot prove that it was easy.

That distinction matters. Bad data should reduce certainty. It should not become a flattering conclusion because the average looks impressively low.

I explain the common metrics and their limits in Training Metrics Made Simple:

Training Metrics Made Simple

My current heart-rate-zone material is here:

my current heart-rate-zone guide

What My Runner Report Added

Before I stared too long at the graphs, I had recorded what the run felt like.

I reported 0 out of 10 pain before, during and after. That was one of the strongest positives from a 50-minute endurance extension.

I also reported that the humidity increased the cost of the workout. Conversation and singing remained available as effort checks, but the final two work sections required more attention. I walked early when I needed to rein in effort, walked cautiously on slick bridges and stopped at traffic lights.

None of that turned the workout into a failure.

The original prescription allowed walking when easy effort stopped feeling easy, when breathing and mechanics no longer agreed with the target, or when safety required it. The workout was designed around controlled endurance, not an uninterrupted pace graph.

Cadence also stayed stable once running. There was no obvious late collapse, and a deliberate form check still produced the expected increase in turnover. That suggested I could organize my mechanics even when the run felt more demanding.

The sensor lost the plot. I didn’t.

Four Sources, Four Different Jobs

Four sources in an AI running coach analysis: prescription, activity recording, runner report and coaching assessment
A coaching assessment becomes more honest when each source is allowed to say only what it knows.

This case became much clearer when I separated the sources.

Coach prescription

The assignment was 50 minutes with four easy-endurance sections at RPE 3 to 4. The purpose was a modest endurance extension, not speed.

Garmin and Strava recording

The devices recorded elapsed time, distance, pace changes, cadence, stops and a corrupted heart-rate stream. Those records tell me what the sensors captured, not why every change occurred.

Runner report

I supplied RPE, breathing and singing checks, humidity, pain, route hazards, traffic stops and the increasing cost of the final sections.

Coaching assessment

CoachChat compared the prescription with the credible evidence, rejected the unusable heart-rate data and decided whether the workout achieved its purpose and what should happen next.

This source ownership is central to my post-run debrief process:

my AI running coach post-run debrief prompt

If an AI blends all four voices into one confident paragraph, it becomes difficult to see where a conclusion came from. If each source is allowed to say only what it knows, uncertainty becomes easier to manage.

What CoachChat Concluded

CoachChat’s verdict was a successful endurance extension with a major heart-rate-data caveat.

The 50-minute duration was completed. Overall RPE matched the prescribed 3 to 4. Cadence and mechanics remained stable. Pain remained at zero. My decisions to slow or walk were appropriate effort control and risk management.

At the same time, the environmental load was substantial, and the later sections carried enough subjective cost that the run did not justify automatically adding another five minutes the following week.

That is not less positive than Strava’s “real improvement.” It is more specific.

The improvement was not “the athlete is now faster in humidity.” The useful achievement was that I extended easy endurance to 50 minutes without pain, kept control when the sensor failed and made sensible adjustments when the conditions raised the cost.

What Happened to the Next Week

The next plan is where an AI running coach analysis becomes a coaching decision rather than a recap.

CoachChat did not punish the run. It did not remove endurance work. It also did not mechanically progress the Sunday session to 55 minutes.

Week 6 held the longest run at 50 minutes so I could consolidate the new duration. Easy running remained governed primarily by RPE and breathing. Heart rate stayed on for observation, but only as a soft validation metric when the chest-strap signal was believable. No intensity was added.

The week’s headline was:

“Consolidate the engine. Trust effort. Let the data observe rather than command.”

That is the smallest-sensible-change rule in action. I introduced the rule in When to Adjust a Running Plan.

The theory was simple. Change only what credible evidence requires, and only for as long as the evidence supports. This run became the concrete test.

The Input Asymmetry Is the Real Story

Comparison of activity-summary inputs with contextual AI running coach inputs
The fairest comparison begins with the information each system received.

It would be easy to frame this as ChatGPT defeating Strava in an AI showdown. That would be catchy and wrong.

Strava had activity metrics and its own history. It produced a quick interpretation from that information.

CoachChat had much more:

The original prescription and purpose

The planned workout structure

Garmin record-level evidence

My runner diary

Weather and terrain

Overall RPE

Pain and symptom reports

Breathing, conversation and singing checks

Walk and stop explanations

A sensor-quality assessment

Previous runs in the training block

Permission to reject data that failed credibility checks

Responsibility for the next prescription

If I had given CoachChat only the same summary metrics, it could have made the same mistakes. A polished answer can’t repair missing context by confidence alone.

That is why this article supports, rather than replaces, my larger ChatGPT running coach experiment:

How I Use ChatGPT as My AI Running Coach at 67

The cornerstone explains the overall system. This run shows why the system needs more than an activity card.

What an AI Running Coach Needs Before It Coaches

Context

What was the workout intended to accomplish? A slow easy run and a failed time trial can contain similar numbers and mean opposite things.

Source ownership

Which facts came from the prescription, the device, the runner and the later assessment? An AI should not quietly turn an inference into a recorded fact.

Chronology

One activity can justify a safety decision. A plan-level change usually needs a sequence of comparable runs. The latest upload should not erase the training block that came before it.

Permission to reject bad data

A number is not trustworthy because it arrived with a decimal place. When heart rate, GPS or another sensor conflicts with the rest of the evidence, the system must be allowed to say, “This metric cannot support that conclusion.”

A decision boundary

The system should say what it can conclude, what it can’t conclude and what evidence would change the next decision.

Without those pieces, an AI may still write a glowing recap, but certainly not contextual coaching.

So, Is Strava Athlete Intelligence Useful?

Yes.

It’s useful for quickly surfacing activity facts, comparisons and possible talking points. It may notice a monthly pace best that I missed. It can make the activity easier to scan and more enjoyable to revisit.

I simply wouldn’t ask it to carry more authority than its inputs can support.

In this case, “fastest average pace of the month” was a reasonable data comparison. “The humidity didn’t slow you,” “you seemed energized,” “two intervals,” and “91 percent recovery zone proves controlled effort” were interpretations that the available evidence could not justify.

Strava gave me a useful summary. CoachChat gave me a contextual assessment because I supplied the missing context.

Neither one was standing beside me on the bridge.

What This N=1 Case Cannot Prove

This is one runner, one workout, one version of Strava Athlete Intelligence and one evolving CoachChat workflow. It doesn’t prove that Strava will misread every activity or that my system will always reach the right decision.

I still find it challenging to classify my RPE. The Garmin and Strava summaries may process moving time, stops and environmental data differently. I don’t have a trustworthy heart-rate record for this session, so I can’t reconstruct the exact cardiovascular cost after the fact.

CoachChat can also sound certain when it shouldn’t. Its advantage here came from the evidence I supplied and the rules we imposed, not from immunity to error.

The fair conclusion is narrower: when two AI systems receive different evidence, their interpretations are not directly comparable. A coaching decision needs more context than an activity summary.

The Runner Still Has to Run the Run

I enjoy the graphs. I enjoy the analysis. Apparently, becoming a running-data nerd was part of my Road to 70.

But this run worked because I could still make decisions without a trustworthy heart-rate number. I used RPE, breathing, conversation, singing, mechanics and common sense on a slick bridge. Then I wrote down what happened before a flattering average could skew my recollection of the run.

That may be the most useful result of this AI running coach analysis.

The watch recorded the run. Strava summarized the recording. CoachChat connected the run to the plan. I ran.

Frequently Asked Questions

Can Strava Athlete Intelligence replace a running coach?

It can summarize and interpret activity data, but it doesn’t automatically know the original workout purpose, every symptom, route interruption, sensor problem or the reasoning behind a longer training plan. Treat it as a useful activity summary, not a substitute for contextual coaching or qualified medical advice.

Why was the average heart rate unusable?

The chest-strap trace contained extensive implausibly low readings and abrupt changes that didn’t agree with pace and cadence. That made the whole-session average and recovery-zone percentage unsuitable for zone compliance, aerobic efficiency or drift claims.

Was the run still successful without trustworthy heart-rate data?

Yes, based on what this particular run was meant to accomplish. I completed the prescribed duration, my overall RPE matched the target, pain stayed at 0 out of 10, cadence remained stable and I used breathing and talk checks to control my effort. That made it a successful workout. It didn’t prove that my fitness had improved.

Why did CoachChat hold the next long run at 50 minutes?

I ran the duration, but the final sections carried substantial environmental and subjective cost. Holding at 50 minutes allowed the new endurance duration to consolidate without adding intensity or pretending the humidity was irrelevant.

Does this prove ChatGPT is a better running coach than Strava?

No. The systems had different inputs and different roles. The lesson is that context, source ownership, chronology and data-quality checks improve the decision. The same AI given incomplete or faulty information can still produce a confident mistake.

AI Disclosure

I developed this article with AI assistance for source checking, structure and drafting. AI is also part of the coaching process being examined. The running experience, diary, opinions and final editorial judgment are mine. I reviewed the original prescription, run assessment and next-week plan so readers can distinguish what was planned, recorded, reported and inferred.

Safety and Privacy Disclaimer

This article describes my personal N=1 experiment. It is not medical advice, diagnosis, rehabilitation guidance or an individualized training prescription. Stop exercise and seek appropriate professional assessment for severe, worsening, unusual or movement-altering symptoms, or whenever you are concerned about your health.

Heart-rate devices can fail. Do not use one questionable activity to diagnose a health problem, revise training zones or overrule symptoms. Activity files may contain routes, timestamps and location information. Review what you share with any online service or AI system.

0 0 votes
Article Rating
Subscribe
Notify of
guest

0 Comments
Oldest
Newest Most Voted
Scroll to Top
0
Would love your thoughts, please comment.x
()
x