The repository provides a seven‑week, hands‑on course that guides learners through building a production‑grade research assistant that automatically fetches arXiv papers, processes them, and answers questions using retrieval‑augmented generation (RAG). Week 1 establishes the infrastructure with Docker Compose, deploying FastAPI (port 8000) for REST endpoints, PostgreSQL 16 (port 5432) for metadata, OpenSearch 2.19 (ports 9200/5601) for search, Apache Airflow 3.0 (port 8080) for workflow orchestration, and Ollama (port 11434) as a local LLM service. Week 2 implements an automated ingestion pipeline via the arXiv API, rate‑limited fetching, Docling‑based PDF parsing, and Airflow DAGs that store paper metadata in PostgreSQL. Week 3 adds production BM25 keyword search in OpenSearch, including index management, query DSL, and relevance metrics. Week 4 introduces intelligent section‑aware chunking, Jina AI embeddings, and hybrid search using reciprocal rank fusion (RRF) to combine keyword and vector results. Week 5 completes the RAG pipeline with Ollama‑hosted LLMs, prompt optimization yielding an 80% reduction (≈6× speedup), Server‑Sent Events streaming, and a Gradio chat interface on port 7861. Week 6 integrates Langfuse for end‑to‑end tracing and Redis caching, reporting 150‑400× latency improvements when caching is effective. Week 7 extends the system to agentic RAG using LangGraph workflows that implement guardrails, document grading, query rewriting, adaptive retrieval, and a Telegram bot for mobile access, while exposing full reasoning steps for transparency. Throughout the course learners work with Python 3.12+, UV package manager, and require ≥8 GB RAM and ≥20 GB disk space.
Read original
github-trending/python