Project 01 / Applied AI · Infrastructure
Doc Insight
A multi-tenant Retrieval-Augmented Generation (RAG) service for document Q&A. Upload text files or PDFs, including scanned ones, and ask questions answered from their content, with parsing and embedding handled asynchronously by a separate worker.
Under the hood
- Async ingestion: the FastAPI API validates and returns 202, while a separate worker parses and embeds documents
- Postgres-backed job queue (SELECT ... FOR UPDATE SKIP LOCKED) with lease-based crash recovery and typed retries
- Text extraction with pypdf, falling back to OCR (PyMuPDF + Tesseract) for scanned pages
- Semantic search over 1536-dimensional OpenAI embeddings with PostgreSQL + pgvector
- Multi-tenant API key auth, with every document and query scoped to its owner
- Per-key rate limits, plus Redis-cached query embeddings and answers
- Prometheus metrics and alerts, OpenTelemetry traces from upload to worker, and logs in Loki
