ProjectsSeptember 29, 2026

Document Q&A with Django and pgvector

Built with
  • Django
  • Django REST Framework
  • PostgreSQL
  • pgvector
  • Gemini 2.5 Flash
Document Q&A with Django and pgvector: illustrated cover

System architecture

Component and data flow diagramCSV Q&A pairs to Gemini embeddings: Batch text. Gemini embeddings to PostgreSQL: Store vectors. Question endpoint to Retrieval service: Embed query. Retrieval service to PostgreSQL: Vector search. Retrieval service to RAG service: Relevant context.INGESTION ABOVE / QUESTION ANSWERING BELOWBatch textStore vectorsEmbed queryVector searchRelevant contextCSV Q&A pairsIngestion serviceGemini embeddings768-dimensional vectorsPostgreSQLpgvector + source dataQuestion endpointQuery + categoryRetrieval serviceCosine similarityRAG serviceGemini / fallback
Swipe horizontally to inspect the diagram.Retrieval applies category and similarity filters. With no relevant context, the service returns a fallback instead of asking the model for an unsupported answer.
A general language model does not automatically know the contents of a private dataset. This backend makes stored Q&A pairs searchable by meaning, then supplies matching text as context for generation. It is a separate project from the multilingual, multi-tenant RAG engine. The ingestion service validates CSV columns, formats question-and-answer rows, requests Gemini embeddings in batches and stores content, source, category and vectors in PostgreSQL through Django models. Failed embedding batches are logged and counted as skipped. At query time, the retrieval service embeds the question using the same embedding path. A pgvector cosine-distance query orders candidates, optionally filters by category and keeps candidates within a relevance threshold. The RAG service builds a prompt from that context and calls Gemini for the answer. If nothing meets the retrieval threshold, it returns a fallback without making the generation call. REST views expose query, ingest and health operations. Model names, dimensions, thresholds and retry settings live in configuration rather than being repeated throughout the services. pgvector keeps content and vectors in the same database. This reduces the number of services to operate, but vector indexing and retrieval quality still need workload-specific measurement. Using a hosted embedding API simplifies local infrastructure while adding external latency, quotas and a data-processing dependency. A relevance threshold helps reject weak context; it cannot guarantee hallucination-free answers. The threshold needs evaluation against representative questions, including questions the dataset cannot answer. The repository documents authentication as unfinished and the Celery task as a disabled stub. Ingestion is therefore described here as a synchronous prototype. Before exposing it publicly I would require authentication, bound upload size, verify safe replacement on partial embedding failures and build a retrieval evaluation set. No production latency or accuracy is claimed for this project.

Related projects

Let’s talk about the engineering

I’m open to software engineering roles across backend, platform, and data teams. Get in touch to discuss the architecture, trade-offs, or how this experience could help your team.
Get in touch