July 1, 2026

Movie discovery

Designed an ML serving workflow combining content similarity, item co-occurrence, and popularity; orchestrated training and exposed recommendations behind a FastAPI contract.

Machine learning project
Role
ML systems and backend engineer
Published
July 2026
Focus
Machine learning project
Engineer
Saaim Abdullah
Recommendation engineering — from raw ratings to a serving API system overview

A recommender is a product system, not a similarity function

A good recommendation model is worthless if a product cannot load it, make repeatable requests, handle a new user, or replace a model without breaking the API. I approached MovieLens as a full ML-systems problem: ingestion, feature preparation, ranking signals, artifact lifecycle, serving, and operational behaviors all belong in the design. I built a recommendation service that combines three complementary signals—content similarity, item co-occurrence, and popularity—and exposes them through FastAPI. Training runs separately from requests, with Airflow coordinating the workflow. Redis caching is optional; model artifacts can be reloaded through an explicit service operation.

Scope in numbers

AreaImplemented architectureEngineering benefit
Ranking signals3 — content, co-occurrence, popularityDifferent evidence for familiar items and cold starts
Main API capabilities4 groups — recommend, similar items, health, model reloadClear serving and operational contracts
Model lifecycles2 — offline training and online servingTraining failures do not block every live request
Source familiesRatings, tags, links, and movie metadataJoin user behavior with item context
Optional cache systems1 — RedisAvoid repeat work where cache keys are safe
SchedulingAirflow workflowRepeatable feature/model preparation
The counts describe components, not measured precision, click-through, or online user activity.

How the system works, from file to response

PhaseEngineering workOutput
IngestRead MovieLens ratings, tags, links, metadataConsistent movie and interaction records
PrepareClean identifiers, merge catalogue attributes, derive featuresContent and behavior inputs
TrainBuild content, co-occurrence, popularity scoring structuresCandidate ranking signals
EvaluateApply offline ranking checks to saved outputsComparable model candidate information
PackageSerialize model and its dependenciesExplicit loadable artifact
ServeFastAPI loads artifact and validates requestRecommendation/similar-item response
OptimizeOptional Redis lookup, structured logs, tracesControlled latency and observable requests
RefreshOperator invokes model reloadChange model independently from API code

Signal 1: content similarity

Content-based ranking uses descriptive movie attributes to find items with comparable characteristics. It remains useful for a user with little interaction history, but it can overconcentrate on attributes already seen and miss collaborative taste patterns.

Signal 2: item co-occurrence

Movies that occur together in user interaction histories can be related even when their metadata differs. This helps capture collective behavior that content features alone miss. The quality depends on sufficient interaction density and on preventing very popular items from dominating the signal.

Signal 3: popularity

Popularity provides a stable fallback when personalization evidence is missing. It is especially valuable for new or anonymous users, but it can overrepresent mainstream items. I treat it as one input into a ranking policy, not as proof that a blended score is automatically better than a simple baseline.

The serving contract matters as much as the model

The API includes recommendation and similar-item responses, health checking, and a model-reload operation. API-key protection creates a basic access boundary. Container configuration, test workflows, structured logging, and request tracing make this more than a notebook experiment. An optional Redis cache can reduce repeated calculations, but the cached result must be tied to the relevant model version and user/context inputs. A reload endpoint alone does not give atomic zero-downtime releases. In a stronger deployment, I would load the candidate into a separate object, validate schema and sample predictions, swap references atomically, and invalidate stale cache entries. Those are operational acceptance criteria, not undocumented claims about the current code.

Trade-offs and what they cost

Engineering choiceWhy it is usefulHidden cost
Blend three signalsCovers content, behavior, and sparse-history casesWeight selection needs offline testing
Artifacts separate from requestsTraining can be reproducible and scheduledVersioning and compatibility become mandatory
FastAPIClear, typed HTTP serving layerRequest concurrency and artifact memory require testing
Optional RedisPotentially reduces repeated ranking workCache invalidation on model changes
Model reload operationModel iteration without full service rewriteThread safety, rollback, deployment coordination
Airflow trainingObservable stage dependenciesScheduler is unnecessary for the smallest local experiments

How I would prove ranking quality

An impressive recommender screenshot is not an evaluation. I would use chronological user-level holdout data to avoid training on future behavior, then compare the hybrid model against popularity-only and each individual signal.
MetricQuestion it answersWhat to report
Recall@10Were relevant held-out items retrieved?Score and eligible-user count
NDCG@10Were relevant items near the top?Score by user segment
CoverageHow much of the catalogue is actually recommended?Percentage of distinct recommended items
Cold-start hit rateHow does the service perform with sparse history?Separate cohorts, not blended overall average
API p50/p95 latencyHow quickly do users receive results?Warm, cold, and cached separately
Reload failure recoveryDoes an invalid artifact disrupt requests?Test outcome, rollback path
@10 represents the proposed evaluation cutoff, not a published result. I would also run ablation tests: remove each signal and measure its marginal contribution instead of assuming all three deserve equal weight.

Result

The implemented project demonstrates a complete recommendation-serving shape: multiple ranking signals, offline preparation, persistent artifacts, API access, reload hooks, and optional caching. It shows how I think about the boundary between a data-science experiment and a service that another application can integrate. It does not claim an unmeasured “30% better recommendations” or commercial deployment. The next improvement is an evaluation report tied to a reproducible dataset split, plus load/reload testing. That would turn the existing implementation into a quantitatively defensible ML case study. For an ingestion and warehouse comparison, see my e-commerce data pipeline.

Source

Recommendation engineering — from raw ratings to a serving API architecture diagram 1

More to explore

Let’s talk

I like working through complex problems with people who care about the details. Have a product to build, an engineering role, or an interesting challenge? Let’s start a conversation.

A little note

SaaimOpen to full-time roles, contract work, and conversations about things worth building.

ϟ 1
Contact