November 15, 2025
Edge ML — splitting an autoencoder between a device and the cloud
Trained a compatible encoder/decoder pair, deployed the encoder toward embedded inference, and connected encoded sensor payloads to a Flask decoder on AWS EC2.
Edge ML prototype
- Role
- Embedded ML and cloud engineer
- Published
- November 2025
- Focus
- Edge ML prototype
- Engineer
- Saaim Abdullah

The design question: what belongs on a constrained device?
Architecture, in five accountable stages
| Stage | Runtime | Input → output | Contract I had to manage |
|---|---|---|---|
| 1. Train | Python / TensorFlow | Sensor features → trained encoder + decoder | Data scaling, model compatibility |
| 2. Export | TFLite / embedded build | Encoder model → device-compatible artifact | Supported operators, tensor shapes |
| 3. Encode | ESP32 / TFLite Micro | Scaled sensor values → latent vector | RAM arena, expected numerical ranges |
| 4. Transmit | Device/network boundary | Latent vector → request payload | Serialization overhead, loss, retries |
| 5. Decode | Flask / EC2 + S3 integration | Latent vector → reconstructed value/artifact | Matching decoder version and inverse preprocessing |
The most important engineering principle: a model is a distributed contract
.tflite file is insufficient: the scaler and decoder are also part of the release.
For example, if the firmware scales an input to [0, 1] but the decoder expects training inputs normalized to [-1, 1], the request can be well-formed while the reconstruction is wrong. The fix is not “more training”; it is a reproducible artifact interface and parity tests across environments.
| Interface property | Device responsibility | Server responsibility |
|---|---|---|
| Scaling | Apply exact training-time transform | Invert the agreed transform |
| Shape | Emit latent vector with correct width | Reject unexpected widths |
| Version | Identify compatible encoder release | Load matching decoder release |
| Numeric format | Serialize values predictably | Parse without implicit shape changes |
| Failure behavior | Retry or retain failed transmissions | Validate and report failed decodes |
Why the model was split
Quantifying compression without inventing a benchmark
compression ratio = original serialized payload bytes / latent serialized payload bytes
For a clearly hypothetical example, if the original payload occupied 100 bytes and the encoded request occupied 40 bytes, then the ratio would be 2.5:1, and the network payload would be 60% smaller. Those are example arithmetic values not results from this device. Device power savings might still be negative if inference consumes more energy than the radio saves.
| Measurement | Required method | Published result |
|---|---|---|
| Reconstruction MAE / RMSE | Compare decoded and original held-out sensor traces | Not established in attached case study |
| On-wire byte ratio | Measure full real JSON/binary payload sizes | Not established |
| Encoder latency | Benchmark repeated inference on ESP32 | Not established |
| Peak memory usage | Record tensor arena + stack + model footprint | Not established |
| Energy per sample | Measure encode and transmit under fixed conditions | Not established |
| Decoder reliability | Inject bad shape, mismatched version, network failures | Not established |
Two implementation decisions I would defend in review
Failure analysis and hardening path
| What goes wrong | Observable effect | Engineering response |
|---|---|---|
| Encoder/decoder versions differ | Reconstructed values drift | Versioned artifact manifest and compatibility check |
| Unsupported TFLite op on firmware | Inference cannot initialize | Operator audit and on-device smoke test |
| Latent payload is malformed | Decoder crashes or returns nonsense | Strict schema/shape validation |
| Wireless connectivity drops | Samples fail to reach server | Bounded local buffering and retry policy |
| Latent serialization is verbose | No transmission benefit | Compare binary and JSON encodings on real messages |
| Sensor characteristics shift | Reconstruction error increases | Drift monitoring and periodic calibration |
Result and why it demonstrates engineering depth
Source



SaaimOpen to full-time roles, contract work, and conversations about things worth building.