November 15, 2025

Edge ML — splitting an autoencoder between a device and the cloud

Trained a compatible encoder/decoder pair, deployed the encoder toward embedded inference, and connected encoded sensor payloads to a Flask decoder on AWS EC2.

Edge ML prototype
Role
Embedded ML and cloud engineer
Published
November 2025
Focus
Edge ML prototype
Engineer
Saaim Abdullah
Edge ML — splitting an autoencoder between a device and the cloud system overview

The design question: what belongs on a constrained device?

Embedded devices are not smaller laptops. Their memory budgets, compute characteristics, wireless links, and firmware deployment model force architectural decisions that a server-side ML engineer can often ignore. I built this project to explore a split autoencoder: encode a sensor reading close to its source and reconstruct that representation on a cloud service. My responsibility spanned both halves of the system: model development in Python/TensorFlow, conversion for TensorFlow Lite Micro, ESP32-side integration, a Flask decoder running on AWS EC2, and S3 integration for encoded payload artifacts. The important achievement is creating a connected edge-to-cloud inference contract—not making an unsupported claim that every autoencoder reduces energy or encrypts data.

Architecture, in five accountable stages

StageRuntimeInput → outputContract I had to manage
1. TrainPython / TensorFlowSensor features → trained encoder + decoderData scaling, model compatibility
2. ExportTFLite / embedded buildEncoder model → device-compatible artifactSupported operators, tensor shapes
3. EncodeESP32 / TFLite MicroScaled sensor values → latent vectorRAM arena, expected numerical ranges
4. TransmitDevice/network boundaryLatent vector → request payloadSerialization overhead, loss, retries
5. DecodeFlask / EC2 + S3 integrationLatent vector → reconstructed value/artifactMatching decoder version and inverse preprocessing
These five stages are verifiable parts of the architectural workflow; they are not five independent production services.

The most important engineering principle: a model is a distributed contract

The Python training process and ESP32 firmware must agree on input order, feature scaling, latent vector width, numeric type, and model version. A silent disagreement can produce plausible-looking but incorrect values. I designed the project around the fact that shipping only an .tflite file is insufficient: the scaler and decoder are also part of the release. For example, if the firmware scales an input to [0, 1] but the decoder expects training inputs normalized to [-1, 1], the request can be well-formed while the reconstruction is wrong. The fix is not “more training”; it is a reproducible artifact interface and parity tests across environments.
Interface propertyDevice responsibilityServer responsibility
ScalingApply exact training-time transformInvert the agreed transform
ShapeEmit latent vector with correct widthReject unexpected widths
VersionIdentify compatible encoder releaseLoad matching decoder release
Numeric formatSerialize values predictablyParse without implicit shape changes
Failure behaviorRetry or retain failed transmissionsValidate and report failed decodes

Why the model was split

Device-side encoding keeps the representation step close to the measurement and avoids hosting the full reconstruction model on a microcontroller. The trade-off is obvious: the device is now dependent on network transport and the backend must deploy a compatible decoder. This is a sensible experiment when the question is whether smaller transferred representations can justify the added complexity. But compressed latent size alone does not settle that question. A fair comparison measures complete serialized payload bytes, inference time, radio use, reconstruction error, and memory pressure on the same workload.

Quantifying compression without inventing a benchmark

I would calculate an on-wire compression ratio as: compression ratio = original serialized payload bytes / latent serialized payload bytes For a clearly hypothetical example, if the original payload occupied 100 bytes and the encoded request occupied 40 bytes, then the ratio would be 2.5:1, and the network payload would be 60% smaller. Those are example arithmetic values not results from this device. Device power savings might still be negative if inference consumes more energy than the radio saves.
MeasurementRequired methodPublished result
Reconstruction MAE / RMSECompare decoded and original held-out sensor tracesNot established in attached case study
On-wire byte ratioMeasure full real JSON/binary payload sizesNot established
Encoder latencyBenchmark repeated inference on ESP32Not established
Peak memory usageRecord tensor arena + stack + model footprintNot established
Energy per sampleMeasure encode and transmit under fixed conditionsNot established
Decoder reliabilityInject bad shape, mismatched version, network failuresNot established

Two implementation decisions I would defend in review

First: use an explicit edge/server split, not hidden remote inference. Even when the server reconstructs the result, the ESP32 must perform meaningful local model work. That exposes real deployment issues—operator support, exported model format, memory allocation, and firmware/model version coupling. Second: separate representation learning from confidentiality. The repository title references encryption, but an autoencoder's latent representation is not cryptographic encryption. If payload secrecy is needed, I would use authenticated transport and appropriately managed encryption keys, independent of model design. Compressibility and secrecy are different properties with different acceptance tests.

Failure analysis and hardening path

What goes wrongObservable effectEngineering response
Encoder/decoder versions differReconstructed values driftVersioned artifact manifest and compatibility check
Unsupported TFLite op on firmwareInference cannot initializeOperator audit and on-device smoke test
Latent payload is malformedDecoder crashes or returns nonsenseStrict schema/shape validation
Wireless connectivity dropsSamples fail to reach serverBounded local buffering and retry policy
Latent serialization is verboseNo transmission benefitCompare binary and JSON encodings on real messages
Sensor characteristics shiftReconstruction error increasesDrift monitoring and periodic calibration

Result and why it demonstrates engineering depth

This is a working architectural demonstration of split ML inference and its cross-platform interfaces: training, export, microcontroller runtime, network payload, cloud reconstruction, and storage integration. It exercises firmware, backend, cloud, and machine-learning integration in one project. The next milestone is a public benchmark with named ESP32 board, compiler/TFLite versions, representative readings, model size, and measured error/bytes/latency/energy. That would establish whether the approach is economically useful for a real sensor deployment; the current project establishes the design and implementation path.

Source

Edge ML — splitting an autoencoder between a device and the cloud architecture diagram 1

More to explore

Let’s talk

I like working through complex problems with people who care about the details. Have a product to build, an engineering role, or an interesting challenge? Let’s start a conversation.

A little note

SaaimOpen to full-time roles, contract work, and conversations about things worth building.

ϟ 1
Contact