Skip to main content

Mean Time to Recovery (MTTR)

Average time to restore service after a production incident.

Mean Time to Recovery (MTTR) measures the average time to restore normal operation after a production incident. It's the operational-resilience metric in the DORA set.


​

What it measures


​

Leanmote has no incident record to read, so it infers recovery from pull requests. A merged PR whose title contains hotfix, revert or HOT-FIX is treated as a recovery; the clock starts at the previous merged PR in that same repository that is not itself a recovery. MTTR is the average of those intervals.


​

How Leanmote calculates it


​

mttr = avg(recovery_pr.merged_at - previous_pr_in_repo.merged_at)


​

  • Recovery is recognised by title only, and the list is short: hotfix, revert, HOT-FIX. A PR titled "fix: null pointer" is counted as a failure by Change Failure Rate but is not a recovery here — the two metrics use different word lists, so their counts don't line up.

  • Matching is on any part of the title, so a PR about "reverting the copy change" also counts.

  • Reported as an average, in hours. There is no median, no percentile selector and no severity labels.

  • The interval is measured per repository, and consecutive recoveries are skipped over: three hotfixes in a row all measure back to the same original pull request, so the second and third read progressively longer.

  • A recovery with no earlier non-recovery merge in its repository is dropped entirely rather than counted as zero.


​

How to interpret it


​

  • Trending down means recovery is getting faster — operationally healthy.

  • Because it's an average with no outlier handling, one pathological interval moves the whole number. Open the drill-down and look at the longest rows before reacting to a jump.

  • The keyword list is the real limit. A recovery merged under a title that doesn't contain one of those three words is invisible here, and a routine pull request that happens to mention reverting is counted as an incident.


​

What to do about it


​

  • Improve observability — alerts that detect incidents earlier shorten the clock.

  • Clarify ownership and escalation paths so the right people are paged immediately.

  • Invest in runbooks for the most common failure modes.

  • Review post-incident: was the recovery time spent diagnosing or fixing? That reveals whether the bottleneck is observability or change-velocity.


​

Related metrics


​

  • Change Failure Rate

  • Deployment Frequency

  • DORA metrics overview

Did this answer your question?