Icon

AABA GP

Step 4: Model Preperation

Step 6: Risk Scoring, Intervention Prioritisation & Business Translation.

To translate the model into an actionable business tool, each order’s predicted probability of failing to secure a driver was converted into a failure risk score and ranked from highest to lowest risk. Where orders had the same risk score, shorter booking lead time was used as a secondary urgency criterion. A Top 10% intervention scenario was then simulated to represent limited incentive or operational capacity. Among the 1,451 highest-risk scored orders, 872 (60.1%) actually failed to secure a driver, compared with an overall failure rate of approximately 36%. This demonstrates that the risk ranking can concentrate higher-risk orders into a smaller priority group, helping GoGoX direct limited interventions toward requests where support is more likely to be needed. (Take note: the analysis identifies suitable candidates for intervention and does not establish that incentives themselves will cause successful fulfilment)

Step 5: Decision Tree (Classification)

A Decision Tree classifier was trained using the 70% training set to predict whether newly placed orders would fail to secure a driver. The model was then applied to the unseen 30% test set, with class probabilities retained as each order’s predicted failure risk. Performance was evaluated using a confusion matrix, precision, recall and F1-score, with particular focus on identifying Failed to Secure orders. The resulting failure probabilities were converted into percentage risk scores and ranked from highest to lowest to support GoGoX’s prioritisation of early intervention.

Decision Tree Interpretation (From Decision Tree Learner): Requested vehicle type formed the first major split in the model, confirming that matching risk differs substantially across vehicle categories. Subsequent splits show that risk also depends on combinations of timing, order price and booking lead time. For example, van orders were generally lower-risk, but van pickups at or before 6:30am recorded a substantially higher failure rate (~58.8%). This demonstrates that high-risk orders are better identified through combinations of characteristics rather than any single factor alone.

Step 3: Exploratory Data Analysis (EDA)

EDA #2: Matching outcome by vehicle type.

Matching risk varies substantially by requested vehicle type. Lorry24 recorded the highest failure rate at approximately 83%, followed by motorcycles (~65%) and sedans (~61%), while vans had the lowest failure rate (~21%). This suggests that requested vehicle type is likely to be an important predictor of whether an order secures a driver.

However, failure rate should be interpreted alongside order volume, as a high-risk vehicle type with few requests may have less overall operational impact than a moderately risky high-demand category.

EDA #1: Checking Target Balance

Around 36% of orders failed to secure a driver, compared with approximately 64% that successfully secured one. Although secured orders form the majority, the failed-to-secure class is still well represented, suggesting the dataset is suitable for classification without an extreme class imbalance.

Modelling implication: A stratified train-test split should be used to preserve this class distribution, and model performance should be assessed using precision, recall and F1-score rather than accuracy alone.

Step 2: Data Preperation

  1. Make sure our target was created correctly.

  2. See whether one class massively outweighs the other.

That matters later when we evaluate the classification model because accuracy alone can be misleading when classes are imbalanced. 

Missing values are changed to $0

Change to date & time

Extracted pickup hour and day of week from Pickup Time to capture potential time-based differences in driver availability and matching risk.

Step 1: Data Cleaning

Negative lead times were treated as invalid and removed. One extreme booking over 1 year was also excluded, while plausible advance bookings were retained. (Data Cleaning ends here)

Pickup hours were grouped into Overnight, Morning, Afternoon, Evening and Night to identify broader time periods associated with matching risk.

Route features: Waypoints were transformed to extract the pickup area, final drop-off area, and number of stops, capturing geographic and route-complexity factors that may influence driver acceptance and matching risk.

Counted secured and failed-to-secure orders to assess class balance before model training.

Converted order counts into percentages to assess whether secured and failed-to-secure orders are sufficiently represented for classification.

Orders were grouped by requested vehicle type and matching outcome to compare how frequently each vehicle category secured or failed to secure a driver.
Answers this question: Do some requested vehicle types have a higher risk of failing to secure a driver?

Calculated the percentage of orders within each vehicle type that failed to secure a driver, allowing matching risk to be compared fairly across vehicle categories.

EDA #3: Matching Risk by Booking Urgency

Matching failure increases as booking urgency rises. ASAP orders recorded the highest failure rate at approximately 42%, compared with around 25% for Same Day and 18% for Advance bookings. This suggests that shorter booking lead times are associated with greater difficulty in securing a driver and may be an important predictor of matching risk.

Modelling implication: Both Lead Time Minutes and booking urgency should be retained as candidate predictors, as earlier booking appears to give GoGoX more opportunity to secure a driver.

Orders were grouped by booking urgency and matching outcome to assess whether shorter booking lead times are associated with greater driver-matching difficulty.
Answers this question: Are more urgent orders more likely to fail to secure a driver?

Calculated the percentage of orders within each booking-urgency category that failed to secure a driver, allowing matching risk to be compared fairly across ASAP, Same Day and Advance requests.

EDA #4: Matching Risk by Pickup Period.

Matching risk varies considerably by pickup period. Overnight orders recorded the highest failure rate at approximately 66%, followed by Night at around 47%, while Afternoon orders had the lowest failure rate at approximately 31%. This suggests that time of pickup is likely to be an important predictor of whether an order secures a driver.

Orders were grouped by pickup period and matching outcome to assess whether certain times of day are associated with greater driver-matching difficulty.
Answers this question: Are certain pickup periods associated with a higher risk of failing to secure a driver?

Calculated the percentage of orders within each pickup period that failed to secure a driver, allowing matching risk to be compared fairly across time periods.

EDA #5: Matching Risk by Pickup Area.

Matching risk varies substantially even among GoGoX’s highest-volume pickup areas. Balestier recorded the highest failure rate at approximately 44%, followed closely by Geylang East (~44%), Yishun (~42%), Punggol (~41%) and Woodlands (~40%). In contrast, Alexandra Hill had a much lower failure rate of approximately 12% despite having the highest order volume. This suggests that pickup location may be an important predictor of matching risk, and high demand alone does not necessarily lead to poorer matching performance.

Orders were grouped by pickup area and matching outcome to compare driver-matching difficulty across locations. High-volume areas were prioritised to avoid over-interpreting failure rates based on very small order volumes.
Answers this question: Which high-volume pickup areas have a higher risk of failing to secure a driver?

Calculated the percentage of orders within each pickup area that failed to secure a driver so geographic matching risk can be compared fairly.

EDA #6: Matching Risk by Route Complexity.

Matching failure does not increase consistently as the number of intermediate stops rises. While orders with 4–5 stops showed relatively higher failure rates, the pattern fluctuates across more complex routes. This suggests that route complexity may not be a strong standalone predictor of matching risk. Failure rates for very high stop counts should also be interpreted cautiously if they are based on relatively few orders.

Orders were grouped by number of intermediate stops and matching outcome to assess whether more complex delivery routes are associated with greater driver-matching difficulty.
Answers this question: Do orders with more intermediate stops have a higher risk of failing to secure a driver?

Calculated the percentage of orders within each stop-count category that failed to secure a driver, allowing matching risk to be compared across different route complexities.

Orders were grouped as ASAP (≤30 min), Same Day (31 min–24 hrs), or Advance (>24 hrs) to capture how booking lead time may affect matching risk.

The dataset was divided into 70% training and 30% testing data using stratified sampling on Matching Outcome, preserving the target-class distribution for fair model evaluation.

45 of 14,551 test orders (~0.3%) returned no predicted class or probability. Removing the pickup and drop-off location features did not resolve the issue, suggesting these fields were not the cause. Given the very small proportion affected, these records were excluded from baseline model evaluation and noted as a limitation for further investigation.

Orders were ranked from highest to lowest predicted failure risk, allowing GoGoX to identify which newly placed requests require attention first.

Orders were first prioritised by predicted failure risk, with shorter booking lead times used as a secondary urgency criterion when multiple orders had the same risk score. Retained the 10% highest-risk scored test orders (1,451 of 14,506) to simulate a limited-capacity scenario where GoGoX targets intervention toward the requests most likely to fail.

Top-10% Intervention Scenario: Of the 1,451 highest-risk scored orders, 872 (60.1%) actually failed to secure a driver, compared with an overall failure rate of approximately 36%. Targeting only the highest-risk 10% therefore concentrates failed orders substantially more effectively than untargeted intervention and captures approximately 16.7% of all observed failures.

Business implication: GoGoX can use the risk ranking to focus limited incentives or operational support on a smaller pool of orders where intervention is more likely to be needed, rather than distributing incentives broadly

Route Complexity
Math Formula
GroupBy
Column Renamer
Final Data
Column Filter
Route Complexity Failure Rate
Math Formula
Target Distribution
GroupBy
Total Orders
Math Formula
Pivot
Column Renamer
Bar Chart
Train Test Split
Table Partitioner
GroupBy
Decision Tree Learner
Column Renamer
Number of Stops
Sorter
Class Balance
Math Formula
Final Model Features
Column Filter
Column Renamer
Vehicle Failure Rate
Math Formula
Column Renamer
Pivot
Column Renamer
Scorer
Risk Prediction
Decision Tree Predictor
Risk Ranking
Sorter
Failure Risk Score
Math Formula
Bar Chart
Risk Ranking
Sorter
Total Orders
Math Formula
Missing Prediction Checking
Row Filter
Pivot
Top 10%
Top k Row Filter
Column Renamer
GroupBy
Rank
Column Renamer
Priority Queue
Sorter
Risk Ranking
Sorter
Urgency Failure Rate
Math Formula
Total Orders
Math Formula
Excel Reader
Checking for errors
Statistics
Duplicate Row Filter
Matching Outcomes
Rule Engine
GroupBy
Column Filter
Missing Value
Which actually failed?
GroupBy
Column Renamer
Pickup Period Failure Rate
Math Formula
Bar Chart
Pivot
Bar Chart
GroupBy
Total Orders
Math Formula
Risk Ranking
Sorter
Column Renamer
Bar Chart
GroupBy
Column Renamer
Date&Time Difference
Pickup Area Failure Rate
Math Formula
Column Renamer
Pivot
String to Date&Time
Booking Urgency
Rule Engine
Top k Row Filter
String to Date&Time
Pickup Timing
Date&Time Part Extractor
Failure Rate
Sorter
Total Orders
Math Formula
Final Lead Time
Row Filter
Drop-off Area
String Replacer
Route Complexity (Separator Count)
String Manipulation
Pickup Period
Rule Engine
Bar Chart
Pickup Area
String Replacer

Nodes

Extensions

Links