Used Car Price Prediction

Gradient-boosted regression model for used-car valuation on large, messy marketplace data. Compared LightGBM and CatBoost, handled mixed categorical features, and balanced model quality against inference speed.

Choose the ZIP for a complete local setup. The notebook-only download requires the datasets and dependencies from that bundle.

Evaluation plan.

The revised workflow selects models on validation data and evaluates the final choice once. Earlier study scores are documented as historical in the download; they are not benchmarks for this revision.

Selection

Validation

choose the model before final testing

Error metric

RMSE

with MAE and R-squared for context

Tradeoff

Runtime

measure fit and prediction time separately

Test split

20%

held out from preprocessing and tuning

What the project tries to solve.

Estimate market value quickly from used-car listing attributes so a customer-facing pricing workflow can return a useful number fast.

Status: Revised study project. Includes data, dependencies, and a smoke-test mode.

Notebook: used_car_price_prediction.ipynb

The ZIP includes the notebook, README, pinned Python dependencies, a command-line runner, and all required CSV datasets.

A useful pricing estimate needs to be reasonably accurate and quick to return. Comparing error with fit and prediction time makes that tradeoff visible.

Vehicle listings combine missing values, inconsistent categories, and wide price ranges. I chose this problem to work through those data-quality issues and compare models on both error and runtime.

How I approached it.

Cleaned and normalized a noisy marketplace dataset with missing values and mixed categorical fields.

Built preprocessing pipelines for numeric and categorical features.

Compared baseline linear and tree models against boosted methods including LightGBM and CatBoost.

Tracked runtime along with RMSE, MAE, and R-squared to judge production usefulness rather than accuracy alone.

What I would improve next.

Add group-aware validation to address potential duplicate vehicle listings.

Export versioned model artifacts and evaluate on a newly collected holdout dataset.

Add error slices by price band and vehicle segment to show where the model is reliable and where it misses.

PythonLightGBMCatBoostScikit-learn

Get in Touch

Send a brief note, or email me directly.

Your name, email, and message go through FormSubmit to my inbox. Please avoid sensitive personal details. Privacy details.

Continue to FormSubmit to complete its spam verification and send your message. If that step fails, use email instead.

Prefer a text? +1 (669) 226-7751