A product that works perfectly when it leaves the factory is not necessarily a reliable product. True reliability is demonstrated by its ability to continue performing as required throughout its intended service life.

Engineers therefore need to understand the actual operating environment and identify the stresses that drive failure.
Rawpixel | Magnific.com

That may sound obvious, but reliability qualification is often under pressure when development teams are trying to control costs or meet a production deadline. The problem is that discovering a reliability issue after mass production has started can be far more expensive than finding it during development.

Reliability therefore needs to be considered from the beginning, starting with the high-level system requirement.

The first question should be: What does the product need to achieve, and for how long?

From there, engineers can establish what constitutes failure and how serious different failures are. Failures may range from minor issues with little effect on operation to major or catastrophic failures that result in system loss or create a safety risk.

Each category needs appropriate acceptance criteria.

Mean Time Between Failures (MTBF) is one commonly used reliability measure, particularly for repairable equipment. However, it is not necessarily the right measure for every application. A vehicle, for example, may have a defined durability requirement expressed in kilometres, while another product may have a required probability of surviving a specified number of operating hours.

The important point is that the requirement needs to be defined before the test programme begins.

The testing itself must then be designed around the conditions the product will experience in service.

This is where reliability engineering becomes more interesting than simply running a product until it breaks. Real products rarely experience constant loads. A vehicle component, for example, may experience relatively modest forces most of the time, interrupted by occasional high loads caused by potholes, braking, acceleration or cornering.

Those peak loads may occur for only a small percentage of the component’s operating life, but they can have a major influence on fatigue.

Engineers therefore need to understand the actual operating environment and identify the stresses that drive failure. Measurements from real-world operation can provide valuable information for developing representative test conditions.

Sample size is also important. One component surviving a test does not demonstrate that an entire production population will meet the required reliability level. Variations in materials, manufacturing processes and assembly can result in different outcomes between individual units.

A properly designed qualification programme therefore requires sufficient samples and appropriate statistical confidence. There is no getting around the fact that reliability qualification can be expensive and time-consuming.

But the alternative is potentially much more expensive. A problem identified during qualification can be corrected before production. A problem discovered after thousands of units have entered the market can become a warranty, recall, reputation and customer problem.

Reliability is therefore not something to be checked at the end of product development. It needs to be designed into the product – and demonstrated before production begins.

For long service-life requirements, however, another question arises: how can engineers realistically test a product for decades of operation or hundreds of thousands of kilometres? The answer lies in accelerated testing.