September 29, 2026

ESP32 temperature inference — proving the model/firmware contract

Connected Python model training and export with a real ESP32 inference path and Flask reconstruction, including a fixed scaler and a two-value latent vector.

Personal prototype
Role
Software engineer
Published
September 2026
Focus
Personal prototype
Engineer
Saaim Abdullah
ESP32 temperature inference — proving the model/firmware contract system overview

Getting a model onto a microcontroller is a systems problem

A TensorFlow model that behaves correctly in Python can fail on an ESP32 for reasons unrelated to predictive quality: incompatible operator sets, memory allocation, different preprocessing, wrong tensor dimensions, or a mismatch between the firmware and server decoder. I built this experiment to make those interfaces explicit using a DHT11 temperature sensor, TensorFlow Lite Micro, and a Flask reconstruction endpoint. This is a focused embedded-inference project, distinct from the broader edge compression case study. Its result is the device-to-server model pipeline, not an assertion that temperature data was efficiently compressed.

What the implementation contains

Engineering factValue in this variantWhy it matters
SensorDHT11Physical device input, not just an offline CSV
Scalar sensor value used1 temperature readingDefines the shape of the model's input
Latent vector width2 valuesDefines the device-to-Flask interface
Model deployment runtimes2 — ESP32 and Python serviceExport/decoder compatibility must be maintained
Decode API/decodeHTTP boundary for reconstructed inference
Firmware model runtimeTensorFlow Lite MicroEmbedded tensor arena and operator registration
Server runtimeFlask with saved decoder/scalerInverse transform and output handling
A single scalar becoming two latent numbers does not demonstrate compression. Before JSON keys and numerical text formatting, it expands the logical count of scalar values from one to two.

End-to-end execution path

StepWhere it happensWhat I engineered
1. Create dataPython training scriptSynthetic temperature examples for repeatable training
2. Scale valuesTraining pipelineEstablish fixed numerical transform expected by both runtimes
3. TrainKerasLearn encoder and compatible decoder weights
4. ExportPython / TFLite toolingGenerate device model file and C header
5. Read inputESP32 firmwareSample temperature from DHT11
6. EncodeTFLite MicroRun model within allocated tensor arena
7. SendESP32 network clientSerialize two-value vector as JSON
8. DecodeFlask /decodeCheck vector shape, predict and inverse-scale temperature
9. RespondFlask APIReturn success response; currently does not expose numeric reconstruction

Training and artifact packaging

The training script constructs synthetic data and fits a scaler to a known temperature range. It exports the encoder for the device while leaving the decoder and scaler available to Python. That establishes three connected artifacts: the embedded encoder, the server decoder, and the normalization definition. A bug in any one can change the semantic meaning of the request.

On-device inference

The firmware reads the sensor, applies the expected scale, invokes TFLite Micro, and serializes the two latent values for transport. TFLite Micro does not manage memory like a general desktop Python process: the tensor arena is explicitly allocated and the operation set must support the exported model. When debugging device inference, I would verify model initialization, tensor types, input/output shapes, and memory use before assuming the network is at fault.

Reconstruction and contract checking

The Flask service loads the matching decoder and scaler, validates the incoming latent shape, reconstructs the scaled value, then applies the inverse transform. Its published response is an acknowledgment rather than the reconstructed number. This is a conscious distinction in the case study: the code demonstrates the internal reconstruction path; publishing the numeric output in the response would require an API change.

The difficult engineering decisions

DecisionReasonRisk addressed
Keep preprocessing constants alignedDifferent scales produce incorrect outputsSilent model drift across environments
Export a compact firmware-compatible modelEmbedded runtime supports fewer operationsModel initialization failure
Use a fixed tensor arenaPredictable bounded memory allocationRuntime memory pressure
Validate a 2-value payloadServer must know the latent shapeInvalid decoder inputs
Separate device encoder and server decoderExperiments with split inferenceFull model need not run on microcontroller
Track artifacts as a single releaseInterdependent versions must matchFirmware/server compatibility errors

How I would test it as an engineering system

  1. Parity tests: run the same input through the Python encoder and the ESP32-compiled model; compare outputs within a documented tolerance.
  2. Schema tests: reject a zero-length, one-value, or three-value latent request when two values are required.
  3. Range tests: feed temperatures at the boundaries of the fitted scaler and beyond its intended range.
  4. Hardware tests: record inference time, tensor arena utilization, wireless retry behavior, and failure recovery on a named ESP32 board.
  5. Model-quality tests: compare original sensor values against reconstructed values and report MAE/RMSE on held-out real readings.
These are a concrete five-part test plan, not five completed benchmark suites.

What the result proves

The repository connects model training, embedded export, ESP32 sensor input, JSON transport, and Flask-side reconstruction. It is evidence that I can work across Python ML tooling, C/C++-adjacent firmware constraints, memory-bound inference, and backend contracts without treating those layers as independent demos. The result does not currently prove power reduction, production uptime, wireless security, or payload compression. The next improvement would be a versioned artifact manifest and an API that returns the reconstructed number so automated end-to-end tests can compare it directly to the sensor input.

Code

ESP32 temperature inference — proving the model/firmware contract architecture diagram 1

More to explore

Let’s talk

I like working through complex problems with people who care about the details. Have a product to build, an engineering role, or an interesting challenge? Let’s start a conversation.

A little note

SaaimOpen to full-time roles, contract work, and conversations about things worth building.

ϟ 1
Contact