MTBF
Mean time between failures (MTBF) is a core reliability metric that measures the average operating time a repairable asset runs between unscheduled breakdowns. Calculated by dividing total operational uptime by the number of failure events, MTBF excludes repair time and planned downtime. Originating in 1950s electronics engineering, it helps maintenance teams evaluate failure frequency, determine preventive maintenance intervals, and justify reliability investments. Together with mean time to repair (MTTR), MTBF directly determines operational equipment availability.
- One failure slab
- This bar is one month on the LT-3 leak tester, 720 hours long. Each crimson slab is a failure. This one starts at hour 330 and lasts 4.0 hours. The slab's width stands for repair time, drawn wider than scale for visibility. Four slabs stop the machine for 10 hours in total. MTBF counts the slabs. It does not add up their width.
- One span between failures
- The indigo segments are running time. This one runs from hour 334, when repair 2 ended, to failure 3 at 480. That is 146 hours. The five segments read 150, 178.5, 146, 158, and 77.5 hours. They add up to 710 hours of uptime. MTBF is that uptime divided by the number of failures.
- The MTBF readout
- Mean time between failures is 710 hours of uptime divided by 4 failures: 177.5 hours. It is an average over the month, not a promise. The actual spans ran from 77.5 to 178.5 hours. A higher MTBF means the machine ran longer, on average, before it stopped. The TPM team tracks it month by month.
- The count of failures
- Four failures this month. Planned maintenance on the seal cartridge removes failure 3 next month. Three failures remain. Uptime rises only from 710 to 712 hours. MTBF jumps from 177.5 to 237.3 hours. The month is still 720 hours. MTBF rose because the divisor fell, not because the machine ran longer.
- MTTR and availability
- The same month gives two more numbers. Mean time to repair is 10 hours of downtime divided by 4 failures: 2.5 hours. Availability is 710 hours of uptime divided by 720: 98.6%. MTBF and MTTR together set availability. Availability equals MTBF divided by MTBF plus MTTR: 177.5 ÷ 180 = 98.6%.
Key facts
- Formula
- Uptime ÷ number of failures
- Asset scope
- Repairable equipment only
- Origin
- 1950s electronics reliability engineering
- Key exclusions
- Planned downtime and repair time
- Related metrics
- MTTR and availability
By Matthew Savas — Founder of Kaizumi. Reviewed 1 September 2026.
Mean time between failures (MTBF) is a reliability metric that measures the average operating time a repairable system or machine delivers between unscheduled breakdowns. It is calculated by dividing total operational running time by the number of failures that occurred during that period. Originating in electronics reliability engineering during the 1950s as a component life estimate, MTBF is now standard across manufacturing, plant engineering, and operational management. The metric applies exclusively to repairable equipment; non-repairable components instead use mean time to failure. MTBF quantifies equipment reliability in time units, allowing maintenance teams to evaluate failure frequency, compare machine performance, and determine preventive maintenance intervals.
Definition and mathematical formulation
MTBF represents the mean duration of normal operating time between breakdown events. It reflects the rate at which unplanned stoppages interrupt production. Computing MTBF requires tracking two variables over a defined evaluation window: the total uptime during which the equipment was actively running, and the total count of unplanned failures during that exact duration.
To calculate MTBF, divide total operational running time by the number of failure events during the period. Total operating time must account only for actual running hours, excluding planned downtime such as scheduled preventive maintenance, shift changeovers, scheduled breaks, and offline periods when production was not scheduled.
MTBF also excludes repair time. The duration required to diagnose, repair, and restart a machine after a failure belongs to mean time to repair (MTTR). MTBF focuses purely on the operating spans between these repair interventions. Consequently, MTBF rises when breakdowns become rarer, not when an evaluation period is lengthened while failure rates remain constant.
Worked calculation for an industrial asset
Consider an LT-3 leak tester operating on a high-volume valve assembly line. The asset is scheduled to run continuously across three shifts, 24 hours a day, over a standard 30-day month, creating 720 hours of scheduled production time.
During the month, the leak tester encounters four separate unplanned breakdowns:
- Failure 1 occurs at operating hour 150 and requires 1.5 hours of repair time.
- Failure 2 occurs at operating hour 330 and requires 4.0 hours of repair time.
- Failure 3 occurs at operating hour 480 and requires 2.0 hours of repair time.
- Failure 4 occurs at operating hour 640 and requires 2.5 hours of repair time.
The total downtime across the month is calculated by summing the individual repair durations: 1.5 + 4.0 + 2.0 + 2.5 = 10 hours. Subtracting total downtime from the 720 scheduled hours yields total uptime: 720 − 10 = 710 hours.
The five continuous operating spans between the start of the month, the four breakdown events, and the end of the month are 150 hours, 178.5 hours, 146 hours, 158 hours, and 77.5 hours. Summed together, these operational spans equal the total uptime of 710 hours.
Dividing the 710 hours of uptime by the four breakdown events produces the performance metrics for the month:
The resulting MTBF of 177.5 hours is a statistical average. No individual running span equaled 177.5 hours during the month; the actual operating spans ranged from a low of 77.5 hours to a high of 178.5 hours.
Relationship with MTTR and availability
MTBF and MTTR jointly govern operational availability, which is the proportion of scheduled time an asset is operational and ready to produce. While MTBF measures operational reliability, MTTR measures maintainability and repair efficiency.
Availability can be computed directly by dividing total uptime by total scheduled time, or by using the relationship between MTBF and MTTR:
Both calculation methods yield an availability of 98.6%.
Availability balances the frequency of failures against the time needed to fix them. For example, an asset with an MTBF of 355 hours and an MTTR of 5 hours achieves the exact same 98.6% availability as the LT-3 leak tester. The second machine stops half as often as the LT-3, but it remains offline twice as long for each repair. Although their availability metrics match, the operational impact on continuous assembly lines can differ significantly depending on whether frequent short stops or infrequent long stops create larger line disturbances.
In lean manufacturing systems, availability forms the first component of overall equipment effectiveness (OEE), along with performance efficiency and quality rate. Teams tracking plant efficiency can evaluate these trade-offs with an OEE calculator or consult comprehensive guides on OEE, the metric TPM moves to see how failure reduction directly lifts total line productivity.
Impact of failure reduction on MTBF
Because MTBF is calculated using the count of failures in the denominator, reducing the number of failures creates non-linear changes in MTBF, even when total uptime remains largely unchanged.
Assume the maintenance team investigates the root causes of downtime on the LT-3 leak tester. They trace failure 3 to a seal cartridge that degrades on a predictable operational schedule. The team implements a planned replacement procedure during scheduled changeovers, eliminating that failure event from operating hours.
In the subsequent 30-day month:
- Total scheduled time remains 720 hours.
- Unplanned failures drop from four to three.
- The 2.0 hours of unplanned downtime from failure 3 is eliminated.
- Total downtime decreases from 10 hours to 8 hours.
- Total uptime increases slightly from 710 hours to 712 hours.
Recalculating MTBF yields 712 divided by 3, which equals 237.3 hours.
Although the machine's uptime increased by only 2 hours (a 0.28% increase), MTBF increased from 177.5 hours to 237.3 hours, an improvement of 33.7%. Conversely, if the asset experiences five failures instead of four in a 710-hour operating window, MTBF drops to 142 hours, representing a 20% decline.
This sensitivity makes MTBF an effective indicator for maintenance programs such as total productive maintenance (TPM) and predictive programs utilizing condition monitoring.
Industrial use cases
Tracking MTBF across operational assets supports capital planning, equipment comparison, and maintenance scheduling.
Equipment comparison
Two machines designed to perform identical tasks can exhibit identical overall availability while having different failure patterns. For example, a machine with an MTBF of 177.5 hours and an MTTR of 2.5 hours stops twice as frequently as a machine with an MTBF of 350 hours and an MTTR of 5 hours. Understanding these differences allows operations managers to assign critical jobs to machines with longer spans between breakdowns, reducing the risk of line-stopping interruptions.
Investment justification
MTBF data provides the empirical baseline needed to justify reliability investments. When engineers consider adding automated lubrication or scheduled component replacements, the financial decision depends on the downtime avoided. Raising MTBF on the LT-3 leak tester from 177.5 hours to 237.3 hours eliminates one 2-hour unplanned stoppage per month. The financial value of the recovered production output can then be directly compared against the procurement and labor cost of the planned maintenance routine.
Industry examples
Manufacturing
A computer numerical control (CNC) milling machine operated for 2,000 hours over a six-month period and recorded 4 unplanned breakdowns. The resulting MTBF was 500 hours (2,000 ÷ 4). The plant introduced autonomous maintenance routines and structured lubrication schedules under a total productive maintenance initiative. Over the subsequent six months, the machine operated for another 2,000 hours but experienced only 2 breakdowns. The MTBF doubled to 1,000 hours (2,000 ÷ 2).
Healthcare
A hospital clinical engineering department monitored a fleet of infusion pumps across an operating window of 8,000 collective run hours in a single year. During that time, the fleet logged 10 unscheduled device faults, resulting in an MTBF of 800 hours (8,000 ÷ 10). The engineering department updated preventive inspection procedures and conducted operator training to address setup errors. In the following year, failures were reduced to 5 across the same 8,000 hours of operation, raising the fleet MTBF to 1,600 hours (8,000 ÷ 5).
Administrative and IT infrastructure
An enterprise database server logged 8,760 hours of uptime over a calendar year while encountering 6 unplanned system outages. The initial MTBF was 1,460 hours (8,760 ÷ 6). After technical teams installed redundant power supplies, upgraded storage hardware, and optimized environmental cooling, the server recorded only 2 unplanned failures across the next 8,760 hours of operation. The MTBF increased to 4,380 hours (8,760 ÷ 2).
Interpretation and common pitfalls
A common misconception in maintenance management is interpreting MTBF as a deterministic schedule or a guaranteed operating lifespan before failure. An MTBF of 177.5 hours does not indicate that the machine will run for exactly 177.5 hours before stopping, nor does it guarantee that failures occur at evenly spaced intervals. As demonstrated in the worked leak tester example, individual running periods varied between 77.5 hours and 178.5 hours. MTBF represents the central tendency of a failure distribution across an operational period.
A second issue occurs when teams confuse MTBF with mean time to failure (MTTF). MTTF applies strictly to non-repairable items, such as incandescent light bulbs, specific solid-state components, or disposable sensors, where the first failure marks the permanent end of the component's operational life. MTBF applies exclusively to repairable systems that are restored to working order after a stoppage.
Finally, accurate MTBF tracking depends on strict separation between scheduled and unscheduled time. If a plant records planned maintenance stops, operator breaks, or no-demand idle time as operational uptime, the calculated MTBF becomes falsely inflated. Conversely, if planned maintenance durations are counted as unscheduled breakdowns, MTBF will be artificially depressed. Reliability teams must ensure consistent time-logging practices across all operating shifts to maintain meaningful MTBF metrics.
Frequently asked questions
- What is the difference between MTBF and MTTF?
- Mean time between failures (MTBF) applies exclusively to repairable equipment that is restored to working order after an unscheduled breakdown. Mean time to failure (MTTF) applies strictly to non-repairable components, such as incandescent light bulbs, disposable sensors, or specific solid-state parts, where the first failure ends the component's operational life. While MTBF measures operational spans between repairs, MTTF measures the expected operating span until permanent disposal.
- Does MTBF guarantee that a machine will operate for that exact duration before failing?
- No, MTBF is a statistical average of operating time between breakdowns, not a guaranteed lifespan or deterministic schedule. Individual running spans can vary widely around the mean; for example, an asset with an MTBF of 177.5 hours may experience individual operating intervals ranging from 77.5 hours to 178.5 hours. It represents the central tendency of a failure distribution across an operational window rather than a fixed interval between failures.
- Why can two machines share the exact same availability while having very different MTBF values?
- Operational availability depends on both failure frequency and repair speed, calculated as MTBF divided by the sum of MTBF and MTTR. An asset with an MTBF of 177.5 hours and an MTTR of 2.5 hours achieves the exact same 98.6% availability as a machine with an MTBF of 355 hours and an MTTR of 5 hours. The second machine stops half as often, but each stoppage keeps it offline twice as long.
- Why does eliminating a single breakdown cause a disproportionately large change in MTBF?
- Because MTBF places the count of failures in the denominator, the metric responds non-linearly to failure reduction. On an asset scheduled for 720 hours, eliminating one breakdown increases total uptime by only 2 hours (a 0.28% increase) while increasing MTBF from 177.5 hours to 237.3 hours, an improvement of 33.7%. Conversely, experiencing five failures instead of four across a 710-hour operating window drops MTBF to 142 hours, representing a 20% decline.
- How does incorrectly recording planned downtime distort the MTBF calculation?
- Accurate MTBF tracking requires strictly separating operational running time from scheduled downtime. If planned maintenance stops, operator breaks, or no-demand idle periods are recorded as operational uptime, the calculated MTBF becomes falsely inflated. Conversely, if scheduled maintenance stops are mistakenly counted as unscheduled breakdowns, the increased denominator artificially depresses MTBF.