Gold Recovery Process Modeling
Multi-stage regression pipeline to predict rougher and final gold recovery from plant telemetry. Cleans process data, validates recovery calculations, tunes models with mean squared error, and evaluates a weighted sMAPE business metric.
Choose the ZIP for a complete local setup. The notebook-only download requires the datasets and dependencies from that bundle.
Evaluation plan.
Hyperparameters are tuned with mean squared error; models are compared on validation weighted sMAPE. Earlier study scores are historical and are not benchmarks for the corrected preprocessing.
2 stages
rougher and final gold recovery
sMAPE
25% rougher and 75% final recovery
MSE
cross-validation objective for tree models
Train only
targets are never imputed
What the project tries to solve.
Predict both rougher-stage and final-stage gold recovery from process telemetry so plant operators can estimate output quality earlier in the pipeline.
Status: Revised study project. Includes data, dependencies, and a smoke-test mode.
Notebook: gold_recovery_process_modeling.ipynb
The ZIP includes the notebook, README, pinned Python dependencies, a command-line runner, and all required CSV datasets.
An early estimate of recovery could support plant monitoring, but only if it uses measurements available at prediction time. The supplied test columns define that constraint in this study.
I chose this problem to connect a two-stage industrial process with feature availability, missing measurements, and a domain-specific evaluation metric.
How I approached it.
Validated the recovery calculation itself before trusting the labels.
Aligned training and test features to avoid relying on unavailable plant measurements at inference time.
Predict rougher and final recovery jointly, then combine their errors with the weighted business metric.
Tune tree models with mean squared error, then compare model families using validation weighted sMAPE.
What I would improve next.
Evaluate a custom weighted-sMAPE scorer for joint hyperparameter tuning.
Add clearer diagnostics around train / test distribution shift and feature drift across process stages.
Explain the business meaning of sMAPE and where the current model would and would not be trusted operationally.