ProjectsSeptember 29, 2026

ESP32 temperature inference experiment

Built with
  • Python
  • TensorFlow
  • TensorFlow Lite Micro
  • ESP32
  • Flask
ESP32 temperature inference experiment: illustrated cover

System architecture

Component and data flow diagramTraining script to Encoder export: Convert encoder. Encoder export to ESP32 inference: Flash model. Decoder export to Flask /decode: Load artifacts. DHT11 sensor to ESP32 inference: Scalar input. ESP32 inference to Flask /decode: HTTP: 2 values.OFFLINE TRAINING ABOVE / DEVICE TO SERVER BELOWConvert encoderFlash modelLoad artifactsScalar inputHTTP: 2 valuesTraining scriptSynthetic temperaturesEncoder exportTFLite / C headerDecoder exportKeras + fitted scalerDHT11 sensorTemperature sampleESP32 inferenceNormalize / encodeFlask /decodeReconstruct / log
Swipe horizontally to inspect the diagram.Training also saves the decoder and scaler for Flask. This experiment maps one scalar to two latent values; it demonstrates split inference, not payload compression.
A trained model is only one part of an embedded ML system. Its input scaling, exported format, device runtime and server-side counterpart all need to agree. This experiment explores those interfaces with temperature readings. The inspected training variant generates synthetic temperature samples, fits a scaler over a fixed range and trains a Keras autoencoder. It exports the encoder as a TensorFlow Lite file and C header, and saves the decoder and scaler separately for Python serving. The ESP32 code reads a DHT11 sensor, scales the temperature using matching constants and invokes TensorFlow Lite Micro. A fixed tensor arena holds the runtime's working memory. The device packages a two-value latent vector as JSON and sends it over Wi-Fi to the Flask /decode endpoint. The server loads the decoder and scaler, checks the incoming vector shape, predicts a scaled temperature and applies the inverse transformation. The inspected endpoint logs the reconstructed temperature and returns a success message rather than the numeric prediction. The model artifacts and preprocessing constants form a contract across Python and firmware. A different scaler, latent dimension or unsupported operation can break that contract even when both programs start successfully. Exported symbol names and registered device operations also need verification together. This particular variant maps one temperature value to two latent values. It demonstrates split inference; it is not evidence of payload compression. It is presented separately from the existing edge-compression case study so their results are not conflated. The repository contains the training, export, firmware and decoder pieces. It does not establish field reliability, energy savings or production deployment. Next I would create a versioned artifact manifest, test Python and device outputs against shared samples, return the reconstructed value explicitly, and measure error on real sensor data. Network retries and authenticated transport would follow before deployment outside a lab.

Related projects

Let’s talk about the engineering

I’m open to software engineering roles across backend, platform, and data teams. Get in touch to discuss the architecture, trade-offs, or how this experience could help your team.
Get in touch