Advanced RAG for Documents

This is a live, interactive demo of the RAG system. Feel free to upload a PDF or an Excel file (even a complex one!) to ask questions.

Project Overview

The Challenge

Standard Retrieval-Augmented Generation (RAG) systems often struggle with real-world business documents, especially Excel files with merged cells, hierarchical headers, and complex layouts. Simply converting rows to text loses critical structural context, leading to inaccurate or incomplete answers.

The Solution

This application solves the problem by providing an intelligent pre-processing layer. Before vectorization, the user can configure how each document is interpreted. This includes strategies like forward-filling to handle merged cells and converting entire tables to Markdown to preserve their 2D structure. This ensures the Large Language Model receives context that is both rich and accurate, dramatically improving the quality of the generated answers.

Key Features

  • Multi-file upload for PDF and Excel (.xlsx).
  • Interactive UI to choose data processing strategies.
  • Intelligently handles merged cells via forward-filling.
  • Converts complex tables to Markdown to preserve structure.
  • Powered by Google's Gemini API for embedding and generation.
  • Deployed on Streamlit Cloud with secure API key management.

Technology Stack

StreamlitPythonGemini APIPandasRAGPyMuPDFNext.jsTailwindCSS