notes from the build

Journal / My Blog

Architecture decisions, things that broke, and what I learned along the way.

NOTE 01

Distributed systems

What Happens When Stripe Charges the Customer but Your Server Crashes?

A successful charge can outlive the server request that initiated it. I walk through a Node.js, PostgreSQL, Stripe, and Kafka design that records orders before payment, publishes events through a transactional outbox, and reconciles payment state through verified webhooks. An inbox makes repeated events safe; delayed retries and dead letter queues handle failures. Saga compensation refunds payments when fulfillment fails. The central lesson: design for repeatable business operations across systems that cannot share a database transaction.

Chapters

  1. A practical guide to reliable distributed payment workflows using idempotency, Kafka, webhooks, transactional outbox, retries, DLQs, and Saga compensation.
  2. Start With a Normal Database Transaction :
  3. Stripe Changes Everything :
  4. The Classic Failure
  5. Idempotency Making Retries Safe :
  6. Don’t Create the Order After Payment:
  7. Initial Order Transaction:
  8. Now We Need Kafka:
  9. The Dual-Write Problem :
  10. Transactional Outbox Pattern :
  11. Payment Service Consumes the Event :
  12. Stripe Web hooks Become the Durable Signal
  13. The Inbox / Processed Events Pattern
  14. Payment Success Transaction
  15. Kafka Delivers PAYMENT_SUCCEEDED
  16. The Hard Failure: Database Commit Succeeds but Kafka ACK Does Not
  17. When Should We Commit the Kafka Offset?
  18. What Happens When Processing Keeps Failing?
  19. Retry With Exponential Backoff
  20. But What If Payment Succeeded and Order Cannot Be Fulfilled?
  21. Saga Pattern
  22. Refund as a Compensating Transaction
  23. Don’t Keep the Original HTTP Request Open
  24. Complete Architecture Diagram / Flowchart :
  25. Failure Cases This Architecture Handles:
  26. Exactly-Once Is Usually the Wrong Goal
  27. Final Mental Model
READ ARTICLE ↗
NOTE 02

Data engineering

Building a Real-Time E-Commerce ETL Pipeline with Kafka, Spark, and Airflow

I follow synthetic order events from a Python producer into Kafka, then use Spark Structured Streaming to write checkpointed Bronze Parquet data. Silver removes duplicate orders and invalid records; Gold builds dimensions and sales facts in PostgreSQL. Airflow schedules the batch transformations while streaming ingestion runs independently. The walkthrough explains local Docker compute, cloud object storage, repeatable writes, and the additional quality checks, deployment changes, and monitoring a production pipeline would need.

Chapters

  1. Why I built this:
  2. Architecture:
  3. Layer 1 Event generation and Kafka:
  4. Layer 2 Spark Structured Streaming into Bronze:
  5. Layer 3 Batch transforms (Bronze → Silver → Gold):
  6. Silver clean and validate:
  7. Gold (star schema):
  8. Layer 4 Orchestration with Airflow:
  9. Local vs Production Comparison:
  10. What I’d add next
  11. Takeaway
READ ARTICLE ↗
NOTE 03

AI architecture

Building a Multilingual, Multi-Tenant RAG Engine in Django

How can a RAG API answer across languages without retrieving another customer's documents? I combine Django authentication with tenant-filtered pgvector searches and local multilingual E5 embeddings. Separate query and passage prefixes keep retrieval aligned with the model's training. Gemini receives only relevant context, with a distance threshold allowing the system to decline unsupported questions. The article connects those boundaries to ingestion, tenant quotas, source attribution, and the need to rebuild embeddings when changing models.

Chapters

  1. Repo:
  2. The problem statement:
  3. Multi-tenancy, explained properly:
  4. Multilingual Within a Tenant Understanding the Difference:
  5. Multilingual Embeddings Running Everything Locally
  6. One Small Detail That Makes a Big Difference
  7. Authentication: tenant identity lives in the token:
  8. This is the difference between a multi-tenant system and a data breach waiting to happen: the boundary is enforced by the token and the query filter, in that order, on every request.
  9. Grounded generation, with an honest “I don’t know”:
  10. The Complete Request Lifecycle
  11. What This Project Demonstrates
  12. Closing
READ ARTICLE ↗
NOTE 04

Engineering lessons

I Built a RAG Back end in Django on PostgreSQL 18 + Windows. Here’s Every Way It Broke

A small CSV-to-answer backend became a debugging lesson in environment files, PostgreSQL schema permissions, installing pgvector on Windows, and mismatched AI providers. I describe moving to Gemini and aligning its embedding output with a 768-dimensional database column before ingesting 61 records. Cosine retrieval and a relevance threshold then ground answers in the stored material. The prototype also exposes its limits: row-based chunking, no reranking, provider quotas, and database privileges that need tightening before production.

Chapters

  1. Everyone says RAG is five easy steps. It was five steps plus a guided tour of every way PostgreSQL can betray you on Windows.
  2. The Plan (naive, doomed)
  3. Boss Fight 1: settings.DATABASES is improperly configured
  4. Boss Fight 2: permission denied for schema public
  5. Boss Fight 3: The Extension That Did Not Exist
  6. Boss Fight 4: The Great API Key Mix-Up
  7. Wait, What Did I Actually Store?
  8. And Then… It Actually Worked
  9. The Best Part: When It Said “I Don’t Know”
  10. Limitations (which I’m calling a “roadmap”)
  11. Should You Build One?
READ ARTICLE ↗

Let’s talk

I like working through complex problems with people who care about the details. Have a product to build, an engineering role, or an interesting challenge? Let’s start a conversation.

A little note

SaaimOpen to full-time roles, contract work, and conversations about things worth building.

ϟ 1
Contact