Gold Recovery Process Modeling

Multi-stage regression pipeline to predict rougher and final gold recovery from plant telemetry. Cleans process data, validates recovery calculations, tunes models with mean squared error, and evaluates a weighted sMAPE business metric.

Choose the ZIP for a complete local setup. The notebook-only download requires the datasets and dependencies from that bundle.

Evaluation plan.

Hyperparameters are tuned with mean squared error; models are compared on validation weighted sMAPE. Earlier study scores are historical and are not benchmarks for the corrected preprocessing.

Targets

2 stages

rougher and final gold recovery

Reporting

sMAPE

25% rougher and 75% final recovery

Tuning

MSE

cross-validation objective for tree models

Preprocessing

Train only

targets are never imputed

What the project tries to solve.

Predict both rougher-stage and final-stage gold recovery from process telemetry so plant operators can estimate output quality earlier in the pipeline.

Status: Revised study project. Includes data, dependencies, and a smoke-test mode.

Notebook: gold_recovery_process_modeling.ipynb

The ZIP includes the notebook, README, pinned Python dependencies, a command-line runner, and all required CSV datasets.

An early estimate of recovery could support plant monitoring, but only if it uses measurements available at prediction time. The supplied test columns define that constraint in this study.

I chose this problem to connect a two-stage industrial process with feature availability, missing measurements, and a domain-specific evaluation metric.

How I approached it.

Validated the recovery calculation itself before trusting the labels.

Aligned training and test features to avoid relying on unavailable plant measurements at inference time.

Predict rougher and final recovery jointly, then combine their errors with the weighted business metric.

Tune tree models with mean squared error, then compare model families using validation weighted sMAPE.

What I would improve next.

Evaluate a custom weighted-sMAPE scorer for joint hyperparameter tuning.

Add clearer diagnostics around train / test distribution shift and feature drift across process stages.

Explain the business meaning of sMAPE and where the current model would and would not be trusted operationally.

PythonPandasScikit-learnsMAPE

Get in Touch

Send a brief note, or email me directly.

Your name, email, and message go through FormSubmit to my inbox. Please avoid sensitive personal details. Privacy details.

Continue to FormSubmit to complete its spam verification and send your message. If that step fails, use email instead.

Prefer a text? +1 (669) 226-7751