ProjectsSeptember 29, 2026

Machine learning and data foundations

Built with
  • Python
  • pandas
  • scikit-learn
Machine learning and data foundations: illustrated cover

System architecture

Component and data flow diagramAdvertising CSV to Train / test split: Features + sales. Train / test split to Linear regression: Training data. Linear regression to Joblib artifact: Serialize. Joblib artifact to Prediction: Load estimator. Prediction to Evaluation: Predicted sales.OFFLINE REGRESSION EXPERIMENTFeatures + salesTraining dataSerializeLoad estimatorPredicted salesAdvertising CSVTV / radio / newspaperTrain / test split20% train / 80% testLinear regressionFit training samplesEvaluationMSE / R² / scatterPredictionHeld-out featuresJoblib artifactSave / reload model
Swipe horizontally to inspect the diagram.Evaluation compares predictions with held-out sales values. This diagram follows the inspected regression exercise; the repository contains other learning notebooks as well.
Training a model is not enough to understand it. Data selection, evaluation splits and serialization affect whether the result means what it appears to mean. These repositories work through those foundations with small, inspectable examples. The Advertising example reads CSV data through pandas. TV, radio and newspaper spending become features and sales becomes the prediction target. A fixed random seed makes the train/test split repeatable. Scikit-learn fits a linear regression, joblib saves the fitted model, and the script reloads it before predicting on held-out rows. Evaluation prints mean squared error and R-squared and plots actual against predicted sales. Separate labs work through cost functions and gradient descent. The Python_learning repository supports this work with examples of array indexing, broadcasting, dataframe joins, grouping and missing-data handling. The inspected regression script explicitly uses 20% of the data for training and 80% for testing. That is the implemented split, even though a reader might expect the reverse. It leaves less data for fitting and is a useful reminder to inspect evaluation code rather than infer a workflow from familiar names. These are learning repositories, not a deployed prediction API. The saved-model round trip makes artifact loading part of the exercise, but the available code does not establish production accuracy or operational readiness. Next I would compare against a simple baseline, use cross-validation to understand split sensitivity, add feature-schema checks and persist model metadata alongside the artifact. Dataset-specific evaluation would come before claiming the model generalizes beyond this exercise.

Related projects

Let’s talk about the engineering

I’m open to software engineering roles across backend, platform, and data teams. Get in touch to discuss the architecture, trade-offs, or how this experience could help your team.
Get in touch