Now live: Innocentive Marketplace. Get shortlisted for funded opportunities today.

Gearbox Remaining Useful Life (RUL) Prediction Challenge

722 Submissions
$19,000 USD
Challenge under evaluation

Challenge overview

OVERVIEW

Accurately predicting how long machinery components will last is central to enabling condition-based maintenance, increasing system uptime, and reducing operational costs. This Coding Challenge invites Solvers to predict the Remaining Useful Life (RUL) of a test gearbox using a real-world dataset that captures the complete degradation cycles of several gear systems under load.

The provided data includes high-fidelity, multi-channel sensor recordings from six full-lifecycle degradation experiments conducted under fixed-speed, fixed-load conditions. Five runs are provided as training data. Five windows of sensor data from the sixth run (test set) must be used to estimate the time remaining before failure at the end of each window. All recordings capture detailed vibration, torque, current, and acoustic measurements at consistent intervals.

The goal is to create an effective and scalable approach to RUL prediction under controlled but realistic conditions. This data-rich environment provides a strong opportunity to develop methods applicable to broader prognostics and health management (PHM) systems in rotating machinery, gear trains, and other mechanical systems.

You are asked to utilize the training data and gear degradation run information to learn about the kinds of signals that can predate failure, using it to inform your model and code to estimate the remaining useful life of the Challenge test run (G_6) at five, predefined timestamps.

Take note: Solvers are warned that the dataset for this Challenge, available for download in the Datasets & Scoring section, is a large zipped file (.zip format) – file size is around 15.3 gigabytes. This zipped file then has a larger storage footprint when unzipped. Please be aware of the file size implications, both on initial download speed, bandwidth/download allowance, and storage on your computer.

Submit your estimated RUL at the five timestamps, your executable code, and a write-up of your approach to compete against other Solvers! The total award pool for this Challenge is $19,000 USD, with a $8,000 USD award for 1st Place and $6,000 shared by up to 5 runners-up. For the top performer across both TII Coding Challenges (including this and the Order Reconstruction Challenge), there is an additional $5,000 USD bonus prize available.

Winning Solvers or teams will also be invited to an exclusive TII event slated for early 2026, with details to follow.

 

Your IP Rights are protected in this Coding Challenge; TII must pay you an award to obtain them.

The best solutions have the opportunity to win $8,000 USD for 1st place and $6,000 shared by up to 5 runners-up for their performance on the leaderboard together with their submission quality. There is a bonus award of $5,000 on offer for the top performer across both TII Coding Challenges. The Challenge requires a written proposal to be submitted, and requisite files meeting the requirements to be uploaded as attachments.

Awards will be contingent upon the theoretical evaluation of the proposal by TII against the Solution Requirements and your ranking on the private leaderboard. To receive an Award, Solvers are required to grant non-exclusive rights to the Intellectual Property (IP) in their proposed solution. There is no assignment of IP Rights with this challenge. Solvers will retain all rights to any proposal not Awarded.

 

Submissions to this Challenge must be received by 11:59 PM (US Eastern Time) on November 28, 2025.

Please review the later Participation Guidance section before submitting a proposal.

Thank you to Xi’an Jiaotong University (XJTU) for providing the datasets used in this Challenge.

- Login or register your interest to start solving! -

 

 

ABOUT THE SEEKER & ELIGIBILITY

Technology Innovation Institute (TII) is a global scientific research center attracting the world’s foremost scientists and researchers. TII leads worldwide advances in artificial intelligence, autonomous robotics, quantum computing, cryptography and quantum communications, directed energy, secure communication, smart devices, advanced materials, and propulsion and space technologies.

TII belongs to the Abu Dhabi Government’s Advanced Technology Research Council (ATRC), which oversees the technology research.

This Challenge opportunity invites innovators, start-ups, research institutes, and university students from anywhere in the world with the skills, resources and knowledge, who are all eligible to participate in the Challenge, except:

  • Employees of the TII and its affiliates; its parent company or other subsidiaries of the parent company;
  • Employees of agents or suppliers of the TII or any of its affiliates, who are professionally connected with the Challenge or its administration;
  • Members of the immediate families or households of the aforementioned;
  • Any person or entity registered or ordinarily resident in a country that is on a sanctions list at any time during this Challenge (including, but not limited to, the Sanctions Lists maintained by the United States, the United Nations and the European Union).

 

THE CHALLENGE

Background

Rotating machinery is at the heart of countless industrial systems – from wind turbines, automotive gearboxes, through to factory automation. Predicting failure in such systems before they occur could dramatically reduce downtime, improve operational efficiency, and prevent catastrophic damage caused by unexpected faults. However, real-world datasets that capture complete degradation cycles (from healthy to failed states) are rare, and often proprietary.

The field of Prognostics and Health Management (PHM) aims to predict the occurrence of failure – either in components or systems – to minimize the risk. Building accurate, predictive models from real-world, physical data would help engineers across TII to save time, reduce waste, and contribute innovative developments in PHM.

The Setup

This Challenge leverages a unique experimental gearbox degradation dataset designed for the study of condition monitoring and Remaining Useful Life (RUL) modeling. The dataset includes six complete degradation runs (G_1 to G_6) recorded on a test rig simulating industrial load conditions. Each run ends in physical failure - including gear tooth breakage, pitting, and spalling - under fixed-speed operation (2400rpm) and two distinct load conditions (17.9Nm and 26.5Nm).

Data was collected from a rich set of sensors, including:

  • Six channels of triaxial vibration, measured on two gearbox shafts and one channel of uniaxial vibration measured on the auxiliary gearbox shaft,
  • A torque sensor, current clamp, and speed sensor,
  • A microphone capturing airborne acoustic emission.

Each sensor channel was sampled at 12.8kHz in 2.56second bursts every 2minutes, totaling 32,768 samples per burst. This time-aligned multi-sensor data enables advanced degradation modelling, feature engineering, and fusion of mechanical signatures.

Figure 1: the Test Rig, including X/Y/Z labels (in red) for the shafts detailed in Table 1.

Figure 2: the data-acquisition system.

The test rig and data-acquisition system are shown in Figure 1 and Figure 2. A drive motor powers the platform, while a load motor applies the required torque. Two fixed-shaft gearboxes serve as transmission stages:

  • Test Gearbox
    • Spur gears (module 1.5, width 15 mm)
    • Gear pairs: 29–95 teeth and 36–90 teeth
  • Auxiliary Gearbox
    • Spur gears (module 2 or 3, width 30 mm)
    • Test Groups 1–3: 55–21 & 88–27 teeth
    • Other groups: 52–24 & 88–27 teeth

Both gearboxes use KHK standard spur gears.

Sensor arrangement:

  • Two triaxial accelerometers on the input-shaft and intermediate shaft bearings of the test gearbox (radial X, Y; axial Z).
  • One uniaxial accelerometer on the auxiliary gearbox intermediate shaft bearing (radial X).
  • Speed sensor on the drive-motor shaft.
  • Current sensor (clamp) on the motor input.
  • Torque sensor on the test-gearbox input shaft.
  • Microphone near the test gearbox.

All channels have been converted to physical-unit amplitudes.

Sampling: 12.8 kHz, 2.56 s bursts every 2 min  32 768 samples per burst.

Load conditions: 17.9 Nm or 26.5 Nm at a fixed 2400 rpm.

Shutdown criterion: intermediate-shaft Y-axis peak-to-peak vibration > 3× baseline.

TII is using this Coding Challenge to test the Innocentive Solver community: can you leverage this dataset and produce a model that can predict the remaining useful life of a gearbox at key points? Your application (consisting of your model, and results/findings/data) will be compared against other Solvers using a regularly-updated leaderboard.


Datasets & Scoring

Solvers will be provided with the full multivariate time-series dataset for degradation runs G_1 through G_5, each annotated with the exact time of failure (in minutes). These runs are to be used as the training set. Note that the Channel 11 sensor data is missing for runs G_2 and G_3.

Please click the following link to begin a download of the dataset: DOWNLOAD HERE

This dataset also includes example submissions for you to see the desired format.

Solvers should note that this file, even zipped, is approximately 15.3gb, and when unzipped is much larger.

The test run, G_6, follows the same format but is provided up to five predefined timestamps, at which Solvers are required to predict the Remaining Useful Life (that is, the number of minutes from each timestamp to the point of failure).

Dataset Introduction

Six full-lifecycle degradation runs (G_1 … G_6), each with eleven channels of time-series data.

Training set: G_1, G_2, G_3, G_4, G_5

Test set: G_6 (five windows leading up to RUL prediction points)

The following tables give additional information about the data columns collected by the data-acquisition system on the Test Rig (Table 1), and the operating conditions of the dataset ID runs that led to gear degradation (Table 2).

Channel

Signal Type (Measurement)

Table 2. Overview of Gear Degradation Runs

Note: We have made changes to the table to clarify that the Speed is the Input Shaft Speed.

In Table 2, you can see that the Test Rig gearbox was run under 2 sets of constant conditions, each resulting in different Remaining Useful Life (RUL). Your task is to predict the remaining useful life of the test run G_6 at five pre-defined timestamps.

 

Use the multi-source time-series data and degradation patterns learned from G_1 through G_5 to inform your methodology/approach and model for prediction.

Each submission to the Challenge will consist of five RUL predictions for these timestamps, and you will be required to give information about the model used. There is no restriction on the modeling approach: Solvers may use signal processing, feature-based models, neural networks, hybrid degradation models, ensemble techniques, or other methods suitable for condition monitoring.

The final score will be calculated using a Relative PHM Asymmetric Score, which penalizes overestimation and underestimation differently to reflect the risk balance in industrial scenarios. The public leaderboard on the Challenge will display your best submission’s relative PHM asymmetric score for two of the prediction points, where the lower the score is (closer to zero), the better your performance. The best predictions for the remaining three points will contribute to the private leaderboard.

You can view the score for every Coding Submission on the 'My Submissions' tab of the Leaderboard.

Details of Metric:


A Relative RMSE will also be used as a tie-breaker metric and submission timestamp will be used as the final tie-breaker (earlier is better).

Solvers will not have access to the timestamps of any of the sensor measurements in the 5 test windows for G_6. However, all six degradation runs follow comparable operational profiles, enabling generalization through data-driven or physics-informed modelling.

 

GETTING STARTED

    1. Download the data. Please note the large size of the zipped file (approximately 15.3gb), which becomes much larger when unzipped. Be sure to double check your connection, bandwidth allowance, and computer storage before downloading!
    2. Begin your exploration of the signals, gear degradation runs explained above, and potential failure modes.
    3. Use your method or approach to estimate the Remaining Useful Life (RUL) at the 5 key timestamps of G_6.
    4. Create and upload your submission – primarily using the attachments. Ensure you use the correct file type (.csv), naming convention, and include methodology description and executable code.
    5. Check the public leaderboard after for your comparative ranking.
    6. Use your score, status, or competition to guide your next submission!

 

SOLUTION REQUIREMENTS

Submissions should predict the Remaining Useful Life (in minutes) at five provided checkpoints in the G_6 test dataset. Each entry must be submitted in a CSV file with four columns:

  • the checkpoint ID,
  • the corresponding RUL prediction,
  • your lower bound (5th percentile) for prediction,
  • your upper bound (95th percentile) for prediction.

Solvers are encouraged to document their methodology clearly and concisely, including any modeling strategies, signal preprocessing steps, and feature extraction techniques. Solutions must be reproducible and based solely on the provided dataset and open-source tools.

You may submit up to 5 submissions per calendar day (UTC), which will be scored and your relative position against other Solvers will be reflected in the leaderboard.

TII is primarily interested in solutions that meet the following must-have requirements:

  • Private Leaderboard score: the lower the score, the higher quality your results. This will be used as the quantitative measure of the Challenge, and is sourced using Asymmetric Relative RUL Score, with Relative Root Mean Squared Error (RMSE) as a secondary metric and tie-breaker.
  • Robust, innovative methodology for providing quality predictions in this Remaining Useful Life (RUL) task
  • Functional and effective code, provided in accessible format
  • Correct format of submission
    • Provide a .csv file with your findings/data/results named G_6_RUL.csv, with four columns: id, prediction, lower bound, upper bound. This must include your predictions for the five checkpoint IDs: test_1, test_2, test_3, test_4, test_5. Upload it to the attachments field.
      • This format is mandatory for running your submission against the leaderboard, so be sure to include id, prediction, lower bound, upper bound values for each of the 5 checkpoint IDs/test sets.
    • Provide executable code, through a link or attachment upload.
    • Submit a concise, written explanation of your methodology, as an attachment.

 

In addition, TII is interested in the following nice-to-have criteria:

  • Robust methodologies and code that can be applied beyond gearboxes: across industries, input metrics/sensor readings, or machine types.


TII will consider both quantitative and qualitative measures to determine success in this Challenge. This includes the private leaderboard score of your predictions, and the qualitative assessment of your methodology, approach, and code quality.

The uncertainty bounds metric (lower and upper bounds, at 5th percentile and 95th percentile) are required, to analyze the reliability of your prediction.

Things to Avoid

Solvers are not allowed to game or ‘hack’ the Challenge. TII welcomes any and all innovative approaches that rely on proven code, and robust methodologies, but will not consider anyone acting outside the spirit of the Challenge.

 

This Coding Challenge has the following features:

  1. Your IP Rights are protected; TII must pay you an award to obtain them.
  2. The best solutions have the opportunity to win $8,000 for 1st place and $6,000 shared by up to 5 runners-up for their performance on the leaderboard together with their submission quality. There is a bonus award of $5,000 on offer for the top performer across both TII Coding Challenges.
  3. The Challenge requires a written proposal to be submitted, and requisite files meeting the requirements to be uploaded as attachments.
  4. Awards will be contingent upon the theoretical evaluation of the proposal by TII against the Solution Requirements and your ranking on the private leaderboard.
  5. To receive an Award, Solvers are required to grant non-exclusive rights to the Intellectual Property (IP) in their proposed solution.

Solvers will retain all rights to any proposal not Awarded.

 

YOUR SUBMISSION

To enter the Challenge and for your submission to be considered for the leaderboard and award, you must submit the following:

  1. Prediction file
    Type: .csv file, uploaded as attachment to the submission form
    Naming convention: G_6_RUL.csv
    Required contents: Your prediction for Remaining Useful Life (RUL), and your lower and upper bounds (uncertainty bounds metric).
    Format

[Please note, the time under prediction in this example format is example data, purely for illustrative purposes. The prediction file example in the downloadable dataset is purposely scrambled, with example purposes, and should not be used to infer information about the remaining useful life.

Your entries for this column must consist of your model’s findings for the predicted RUL at each timestamp, expressed as a positive number in minute count. For example, if the remaining life was 2 hours, this would be expressed in the number of minutes, and written in the column as “120”.]

Details:

  • Column titled id must match the five prediction timestamps (test_1…test_5).
  • Column titled prediction is your predicted time (in minutes) from each timestamp until failure, known as the ‘remaining useful life’ or RUL.
  • Column titled lower_bound should be lower bound of a (two-sided) 90% confidence interval for your prediction
  • Column titled upper_bound should be upper bound of a (two-sided) 90% confidence interval for your prediction.

 

  1. Executable Code:
    Type: GitHub repository, library, distribution package uploaded, Jupyter notebook as links in relevant form field, or uploaded as attachment to the submission form.
  2. Methodology description:
    Type: Uploaded in PDF or text file format as an attachment to the submission form.
    Details: Written answer describing how you developed your solution, your modelling approach, any difference from previous approaches, and any assumptions or observations.
  3. You will also be asked about your Participation Type (Individual or Organization), and your relevant Experience in the submission form. Please ensure you answer the Experience point in at least one of your submissions.
     

Remember, you can submit up to 5 submissions per day (UTC), with these being reflected regularly on the public leaderboard so you can track your performance.

 

PARTICIPATION GUIDANCE

  1. Submission Close Date: Submissions to this Challenge must be received by 11:59 PM (US Eastern Time) on November 28, 2025.
  2. Late submissions: Late submissions will not be considered.
  3. Multiple submissions allowed: In this Challenge, multiple submissions by the same Solver/team/organization are encouraged: as you optimize your model, the findings/data should improve and you can resubmit to the Challenge. You can check your performance against the regularly-updated public leaderboard on the Challenge.

    Your best-scoring sub
    missions of the Challenge will determine your final position on the public and private leaderboard. Ties will be resolved using the relative RMSE score.
  4. Evaluation notification steps: After the Challenge submission close date, TII will review and select the winning proposals according to the timeline in the Challenge header. Everyone who submits a proposal will be notified about the status of their submissions.
  5. Use of AI: Please note that any submissions produced solely with generative AI are not of interest.
  6. Learn more: Find out more about participation in Innocentive Challenges.

 

Interested in Coding Challenges? We’re also running another opportunity with TII in parallel – the Order Reconstruction Coding Challenge.

We’re looking for Solvers who can find the real running order of 50 files worth of time-series sensor data from vibration recordings. Interested? Join the Order Reconstruction Coding Challenge today!

Use the slider to explore how the Challenge process works:

Register

Review & Accept

Submit

Win

To start solving this Challenge, log in to the Challenge Center or register as a Solver