September 29, 2026

ML foundations — reproducible experiments before model complexity

Built inspectable pandas/scikit-learn regression workflows, trained and reloaded a persisted model, and examined how evaluation splits affect claims of model quality.

Learning project
Role
Software engineer
Published
September 2026
Focus
Learning project
Engineer
Saaim Abdullah
ML foundations — reproducible experiments before model complexity system overview

Knowing the algorithm is not enough; the experiment must be correct

I built these machine learning exercises to understand what a trained model actually means. It is easy to call .fit() and print a score. It is harder—and much more useful—to explain what data the model saw, how inputs were shaped, whether the score is comparable across experiments, and whether an artifact behaves the same way after it is loaded again. This body of work is foundational engineering practice, not a deployed commercial ML platform. I include it because the discipline behind evaluation, preprocessing, and reproducibility is the same discipline required in larger data and machine learning systems.

Quantified scope

ComponentExact structure in the exerciseEngineering significance
Regression input features3 — TV, radio, newspaper spendingMulti-input feature matrix
Regression target1 — salesSupervised continuous prediction
Split recorded in script20% train / 80% testImplementation fact, not an assumed standard 80/20 split
Named evaluation measures2 — MSE and R²Error magnitude and explained variance
Artifact lifecycle2 steps — save and reload via joblibTest that inference can run from a persisted model
Scientific stackpandas, NumPy, scikit-learnData manipulation, numerical computation, ML pipeline

From raw CSV to an evaluated prediction

StageImplementationWhat I check
LoadRead advertising CSV with pandasColumn names, row counts, nulls, data types
SelectBuild X from 3 spending columns and y from salesNo target leakage into predictors
SplitUse deterministic random seed and test/train allocationRepeatability and proper holdout separation
TrainFit linear regression with scikit-learnFeature shape and coefficient stability
SaveSerialize model with joblibArtifact survives process boundaries
ReloadLoad saved estimatorPredictions use saved parameters, not a new fit
EvaluateCompute MSE and R²; actual-vs-predicted plotError and systematic residual patterns

Why I call out the unusual split

The implementation uses 20% for training and 80% for testing. A reader might assume it is the opposite. I do not silently “correct” that fact in my project story. Fewer training rows can make coefficient estimates less stable, especially as model complexity rises. It can also make evaluation results sensitive to the particular sampled records. A mature analysis would rerun the experiment with multiple seeds, compare allocation strategies, and report distributions rather than highlighting one lucky metric.

Interpreting the numbers instead of merely printing them

Mean squared error (MSE) is the average squared prediction error; it is affected by the target's units and penalizes larger misses disproportionately. R² measures relative explanatory performance compared with predicting the mean, and can be negative on a held-out set. Neither alone establishes business usefulness, causal effect, or out-of-sample robustness.
MetricMathematical ideaFailure mode it helps reveal
MSEAverage of (actual - predicted)²Occasional very large misses
R²1 - SS_res / SS_totModel not beating a simple mean baseline
Residual visualizationInspect error across observed salesTrends or nonlinearity ignored by linear model
I would also add MAE and confidence intervals in a future report—but I do not claim that those are already published results.

The principles behind the smaller Python labs

The related Python and data-library exercises cover array shapes, broadcasting, dataframe joins, grouping, and missing-value workflows. Those concepts sound elementary until a production pipeline silently joins records on the wrong key or broadcasts a vector across the wrong axis. In my larger data projects, the same fundamentals determine whether a feature pipeline returns trustworthy records. The separate gradient-descent and cost-function exercises clarify optimization rather than relying solely on a library API. My goal was to understand where model parameters come from, what objective the training process minimizes, and what numerical behavior means when the optimization path diverges or converges.

Evaluation and deployment boundaries

What is establishedWhat still needs evidence
A reproducible training/test split existsSensitivity across multiple random seeds
A linear regression model is trainedBetter-than-baseline behavior across datasets
A model artifact can be saved and loadedInput schema/version metadata embedded in the artifact
MSE and R² are computedRepeated cross-validation and uncertainty ranges
Basic data operations are exercisedAutomated pipeline and production inference service

Result

The deliverable is a set of readable, reproducible ML and data-processing exercises that make choices visible rather than hiding them behind a high-level notebook. Their value is technical precision: controlled splits, explicit features, model persistence, and metrics that can be interpreted critically. Next milestone: build a scikit-learn Pipeline that owns preprocessing and prediction together, add schema validation, compare against a DummyRegressor, and run repeated validation with recorded seeds. That would make this foundation a stronger quantitative case study without pretending it is already a production ML service.

Code and experiments

ML foundations — reproducible experiments before model complexity architecture diagram 1

More to explore

Let’s talk

I like working through complex problems with people who care about the details. Have a product to build, an engineering role, or an interesting challenge? Let’s start a conversation.

A little note

SaaimOpen to full-time roles, contract work, and conversations about things worth building.

ϟ 1
Contact