September 29, 2026

Document Q&A — building a retrieval system that knows when not to answer

Built Django document Q&A with configured 768-dimensional embeddings, top-5 retrieval, batch imports, category filtering, retry handling, and a no-context fallback.

Personal prototype
Role
Software engineer
Published
September 2026
Focus
Personal prototype
Engineer
Saaim Abdullah
Document Q&A — building a retrieval system that knows when not to answer system overview

The product problem: answer from the dataset, not from a guess

A language model can write an excellent-sounding response while knowing nothing about a company's private documents. I built this backend to make answer generation depend on retrieved source records. The service accepts structured Q&A material, turns it into searchable embeddings, selects relevant records for a question, and then asks Gemini to answer from the selected context. If the search finds nothing credible, the request returns a fallback rather than paying for an unsupported generated answer. The project demonstrates a practical backend-engineering point: the most important part of RAG is not the prompt. It is the data contract, retrieval boundary, dependency failure handling, and decision to decline weak answers.

System facts and configured numbers

Metric or configurationDocumented valueInterpretation
Public API operations3 — query, ingest, healthSeparate serving, data administration, and readiness paths
Embedding width768 dimensionsVector schema and embedding model must agree
Candidate limitTop 5Bound the context admitted to generation by default
Cosine distance threshold0.7Filter candidates with distance above the configured cutoff
Ingestion batch size20 recordsBound each embedding request during CSV import
API quota retry setting3 attemptsRecover from transient provider rate limits with backoff
Required input columns4 — name, category, question, answerExplicit and testable ingestion contract
Generation serviceGeminiExternal inference dependency, distinct from PostgreSQL
These values describe repository configuration, not an accuracy score or load-test result. The source README documents them; deployments can tune them.

End-to-end request and data flow

StepInputWork performedWhat can fail
1. ImportCSV fileValidate required columns and normalize Q&A rowsSchema errors, malformed content
2. EmbedText batchesRequest Gemini embeddingsQuotas, model changes, network errors
3. PersistContent + metadata + vectorsStore in PostgreSQL with pgvectorPartial imports, incompatible dimensions
4. QueryUser question + optional categoryEmbed query with same model familyProvider latency, unavailable API
5. RetrieveQuery vectorRank stored vectors by cosine distance; filter by categoryPoor matches, unbounded search cost
6. GateRetrieved candidatesApply top-K and distance cutoffRelevant evidence absent
7. GenerateSelected Q&A recordsBuild constrained prompt and call GeminiUnsupported assertions, timeout
8. RespondAnswer + contextReturn answer and source context, or fallbackClient contract, sensitive source exposure

Import is part of the product, not a one-off script

The same ingestion service can be reached through the web API and management command. It validates the required schema before embedding, processes the input in 20-record batches, and stores the source/category metadata needed by retrieval. The repository implements replacement semantics for the same source filename, reducing duplicate accumulation on repeated imports. Replacement has a subtle correctness risk: deleting an old source before an entire replacement succeeds can lose previously searchable records. The published code should not be represented as a fully transactional, atomic blue/green ingestion pipeline. A stronger next version would stage a source revision, confirm row and embedding counts, then switch the active revision in one database transaction.

Retrieval has a refusal boundary

For a query, Django first requests an embedding, then uses pgvector cosine distance to rank nearby rows. The optional category filter constrains the candidate pool. A distance threshold and top-K cap bound how much evidence enters the prompt; if no row passes the threshold, the system follows a defined no-context path. The distance cutoff is not a probability of correctness. A 0.7 cutoff is a retrieval parameter that should be calibrated against question sets and expected false positives. I would separately measure relevance of retrieved records and faithfulness of the generated answer; good search results do not automatically guarantee good generation.

The API is deliberately smaller than the architecture

EndpointRoleContract
POST /query/Query a stored knowledge baseQuestion, optional category; answer and retrieved context
POST /ingest/Import Q&A CSVValidates columns and reports ingested/skipped records
GET /health/Health inspectionIncludes database connectivity and stored chunk count
Thin API views call reusable services for embedding, ingestion, retrieval, and prompt construction. That division matters because a management command, synchronous request, and future worker should not each implement different versions of the same data rules.

The engineering trade-offs

ChoiceWhy I made itCost or limitation
PostgreSQL + pgvectorOne database stores content, source metadata, and vectorsIndex strategy matters as corpus grows
Hosted embeddingsSimplifies local model operations and deploymentAdds provider latency, quotas, and data boundary
Top-K + thresholdKeeps generation grounded in bounded evidenceNeeds dataset-specific relevance calibration
Source/category metadataMakes results inspectable and filterableImport mapping must be consistent
Provider retry logicHandles transient 429 errorsRetries need budgets and can increase latency
Refusal without retrieved contextReduces unsupported, costly generation callsMay decline answerable questions if retrieval misses

A concrete operational failure scenario

Imagine a 20-row batch where the embedding provider succeeds for some work but the network fails before the import finishes. “Try the whole CSV again” sounds straightforward until you consider duplicated rows or prematurely deleted previous content. I designed the code with source-level re-import behavior and explicit skipped counts, while recognizing that correctness under partial replacement requires staged revisions and integration tests. This is the difference between a happy-path demo and a system one can reason about under failure.

What is implemented, and what is not

Implemented in the repository: Django/DRF entry points, Q&A CSV ingestion, 768-dimensional vector storage, source/category metadata, cosine ranking, configured relevance cutoff, Gemini-based generation, quota backoff, and health output. Not presented as completed production hardening: required authentication, safe limits for uploaded data, atomic source replacement, a running Celery ingestion worker, measured latency/accuracy, or a deployed high-scale benchmark. The Celery task is a disabled stub. Those are clear follow-up tasks, not achievements I am claiming.

Result and measurable next milestone

The delivered result is an end-to-end, runnable retrieval-and-generation service with explicit configuration and inspectable source context. The most valuable behavioral result is deterministic handling when retrieval finds no acceptable evidence: it can return a fallback without a generation call. I would evaluate it with at least four test groups: answerable in-category questions, semantically related but unanswerable questions, category mismatches, and provider failures. Report Recall@5, context precision, answer faithfulness, p50/p95 latency, and failed import recovery against a named dataset. Until those evaluations run, the project claims the implementation—not invented success percentages. For a broader ownership boundary, see my multilingual, multi-tenant RAG engine.

Source and technical walkthrough

Document Q&A — building a retrieval system that knows when not to answer architecture diagram 1

More to explore

Jul 15, 2026
Jul 15, 2026

Fitter Health

Sole engineer across a 106-table healthcare platform, 135 passing tests, AWS delivery, and durable notification recovery.

View project
Client deliveryFull stack
System overview for Fitter Health

Let’s talk

I like working through complex problems with people who care about the details. Have a product to build, an engineering role, or an interesting challenge? Let’s start a conversation.

A little note

SaaimOpen to full-time roles, contract work, and conversations about things worth building.

ϟ 1
Contact