Accurate prediction of a system’s remaining useful life (RUL) is critical for maintenance planning and cost-effective operations. However, real run-to-failure data are usually scarce, which limits model training. Synthetic data generation through augmentation is a promising solution that can enrich training sets by creating realistic failure examples. This thesis investigates what is the optimal ratio of real and augmented data to improve RUL model performance. To achieve this, a series of experiments was conducted using LSTM-based models for the evaluation on the NASA C-MAPSS turbofan engine datasets (FD001 and FD004), and for the augmentation Gaussian Noise Injection and Conditional Generative Adversarial Networks (cGANs) were used.
The results show that the optimal data mix depends on dataset complexity. For the simpler FD001 set, adding augmented data to the original dataset until the synthetic data represents 40–45% of the whole dataset yielded the lowest prediction error. It is not as clear for the more complex FD004 set, where the best mixture is somewhere in a wider span of roughly 10–50% synthetic data, with no single ratio clearly dominating. Importantly, the augmentation method (Gaussian vs. cGAN) had little effect on this optimal range. These and additional findings provide a good starting point for balancing real and augmented data in RUL tasks, however, further research is certainly needed to confirm these trends in other scena