May 10, 2026

Multilingual RAG

Engineered a Django/PostgreSQL knowledge service that separates customer datasets, embeds multilingual documents locally, and constrains generated answers to authorized retrieval context.

RAG and AICustomer data isolation
Role
Backend and ML systems engineer
Published
May 2026
Focus
RAG and AI / Customer data isolation
Engineer
Saaim Abdullah
Multilingual, multi-tenant RAG — isolation before intelligence system overview

The risk is not a bad answer. It is the wrong customer's answer.

A multilingual knowledge assistant becomes a different engineering problem when multiple customers share its infrastructure. Incorrect retrieval might return an unhelpful paragraph; incorrect tenant isolation might return another company's private document. I designed this project around that distinction: every stage from upload to vector similarity to answer generation needs an unambiguous ownership boundary. The outcome is a Django-backed retrieval system with local multilingual embeddings, PostgreSQL/pgvector storage, authenticated customer-scoped lookup, request quotas, and Gemini-based responses. It covers six distinct text workflows while keeping source ownership part of the data model rather than an optional query hint.

Delivery and scope in numbers

Design dimensionImplemented scopeWhat it demonstrates
Supported prompt workflows6Q&A, summary, translation, glossary, extraction, explanation
Knowledge ownershipTenant-associated documents and chunksQuery scope belongs to authenticated identity
Embedding provider type1 local multilingual E5 pathSource text can be vectorized without an embedding API request
Generation provider type1 external Gemini pathOnly selected context is passed to the generation dependency
Principal security gates3 — identity, authorization, retrieval scopeEach protects a different part of the lifecycle
PersistencePostgreSQL plus pgvectorMetadata and vectors share a relational model
Usage controlPer-tenant quotasLimit consumption independently from data isolation
No client count, measured query latency, accuracy, or cross-language benchmark was supplied with the project, so none is invented here.

End-to-end architecture

StageInput and operationRequired invariant
1. AuthenticateVerify token and resolve caller identityNever trust a tenant ID supplied merely as user input
2. Authorize ingestionAccept document for permitted customerCaller must have rights to write that knowledge base
3. Extract and chunkConvert supported content to passagesPreserve source and owner for every chunk
4. Embed locallyApply multilingual E5 encodingUse compatible preprocessing for documents and queries
5. PersistWrite content, metadata, vectors and tenant relationshipTenant ownership survives all data transformations
6. Check quotaCount/limit permitted requestsQuota decisions do not bypass authorization
7. RetrieveSearch only authorized vectors in pgvectorFilter by tenant inside retrieval, not after results are returned
8. GenerateGive relevant passages to GeminiBound external context to authorized retrieved material
9. ReturnRender answer and supporting contextUser receives only their own accessible information

Why PostgreSQL is central to the architecture

I kept source records, chunk metadata, customer ownership, and vectors together in PostgreSQL. That allows the query to combine ordinary relational predicates with vector similarity. The tenant predicate has to constrain the candidate set before results are used for answer generation; it cannot be a cosmetic filter added to the client response. This also simplifies application ownership: there is one place to understand the relationship between a source document and the vectors created from it. It does not remove the need to tune indexes, test expensive searches, or audit every data-access path.

Why the embedding and generation models are deliberately separate

E5 runs locally for multilingual embeddings; Gemini handles response generation from retrieved passages. These are different trust and operating boundaries. Local embeddings reduce dependency on a third-party embedding API for each document and question. Generation still sends selected context to an external model provider. A production privacy review would decide which customers and document classifications permit that transfer. The boundary matters for performance as well: embedding cost, database lookup, and generation latency contribute differently to an end-to-end request. I would trace those stages separately rather than reporting one opaque “AI response time.”

Six capabilities, one retrieval discipline

WorkflowExample useKey quality check
Question answeringAsk a factual question about uploaded contentAnswer is supported by retrieved passage
SummarizationSummarize material in a customer's knowledge baseImportant facts are preserved
TranslationTransform text across supported languagesMeaning is retained, not invented
Glossary creationExplain domain terms from documentsTerms trace back to permitted material
Information extractionPull named fields or factsMissing fields are handled explicitly
ExplanationClarify technical source passagesReadability improves without fabricating claims
Six templates are a product capability, not evidence that all six tasks reached a particular BLEU, recall, or answer-quality score.

Threat model and design reasoning

Failure modeImpactArchitectural response or required verification
Forged tenant_id in requestCross-customer exposureResolve tenant from verified server-side identity
Ingestion under wrong ownerPermanent data misclassificationAuthorize writes and store immutable provenance
Vector search omits owner filterForeign chunks enter promptAdd negative cross-tenant tests on every retrieval path
Email domain assumed to imply membershipUnauthorized tenant assignmentVerify organization membership independently
Concurrent requests race a quota counterUsage exceeds intended limitUse atomic accounting and concurrency tests
Prompt injection inside a retrieved documentModel follows untrusted content as instructionsSeparate system policy from retrieved source content
External generation service receives sensitive textUnintended data egressApply explicit provider/data classification policy
Isolation and quotas solve different problems. A customer with no quota remaining should not query. A customer with plenty of quota must still be unable to query someone else's information. Treating these as the same middleware check would be a security design mistake.

Key trade-offs

I chose a familiar Django relational service rather than splitting ingestion, retrieval, billing, and authentication into new services. At this scale, consistent identity propagation matters more than multiplying deployments. PostgreSQL plus pgvector trades a specialized vector-only database for simpler transactional association with users and documents. Local E5 saves hosted embedding calls but creates a CPU/memory deployment responsibility. Every trade-off is visible in the operating model.

What the implementation establishes

The project joins authenticated retrieval, customer-linked data, multilingual embedding, six prompt workflows, and usage limits into one API-backed product prototype. The article and repository provide the architectural walkthrough. The public material does not establish a commercial deployment, a named number of customers, a measured zero-leak result, or a published multilingual accuracy score. Acceptance criteria for the next hardening iteration: 100% rejection in a deliberately constructed cross-tenant access test suite; documented retrieval Recall@K by language; bounded per-tenant concurrency; recorded p50/p95 stage latency; and a reviewed data-flow policy for the generation provider. These are proposed targets and tests, not observed achievements. Compare this tenant-aware architecture with my smaller document Q&A backend, which focuses on ingestion, configured embeddings, retrieval, and a refusal boundary.

Explore the work

Multilingual, multi-tenant RAG — isolation before intelligence architecture diagram 1Multilingual, multi-tenant RAG — isolation before intelligence architecture diagram 2

More to explore

Let’s talk

I like working through complex problems with people who care about the details. Have a product to build, an engineering role, or an interesting challenge? Let’s start a conversation.

A little note

SaaimOpen to full-time roles, contract work, and conversations about things worth building.

ϟ 1
Contact