A Flask web application for comparing five distinct retrieval strategies using a shared OpenAI-compatible LLM backend. Features a minimal white & coffee design with serif typography.
- Mode 1: Traditional RAG - Vector similarity search using ChromaDB with local embeddings
- Mode 2: Agentic RAG - LLM-driven search with reasoning and expansion
- Mode 3: Vectorless RAG - PageIndex tree structure with human-like navigation
- Mode 4: Graph RAG - Knowledge graph with entity extraction and graph traversal
- Mode 5: Hybrid Keyword RAG - BM25 + semantic search with cross-encoder re-ranking
- Minimal Design - Clean white background with coffee (#6F4E37) accents
- Serif Typography - Georgia/Times New Roman throughout
- Fixed Layout - Viewport-locked design with independent scrolling
- Markdown Rendering - Assistant responses formatted as rich Markdown
- Processing Indicators - Animated status messages during AI processing
- Real-time Metrics - Token usage, retrieval time, and latency tracking
- Python 3.10+
- pip
-
Clone the repository
git clone <repository-url> cd RAGPro
-
Install Dependencies
pip install -r requirements.txt
For local embeddings, this will also download the sentence-transformer model (~80MB) on first use.
-
Configure Environment
cp .env.example .env # Edit .env with your API credentials -
Run the Application
python app.py
-
Open Browser Navigate to
http://localhost:5000
Create a .env file in the project root:
# OpenAI-compatible API Configuration
OPENAI_API_BASE=https://api.openai.com/v1
OPENAI_API_KEY=your_openai_api_key_here
# Model Configuration
MODEL_ID=gpt-4o-mini- OpenAI - Official OpenAI API
- Ollama - Local LLMs (
http://localhost:11434/v1) - Groq - Fast inference API
- Any OpenAI-compatible API
RAGPro/
├── app.py # Flask application entry point
├── requirements.txt # Python dependencies
├── .env.example # Environment template
├── .gitignore # Git ignore rules
├── README.md # This file
│
├── engine/ # RAG engine implementations
│ ├── traditional.py # Mode 1: ChromaDB with local embeddings
│ ├── agentic.py # Mode 2: Agentic RAG with LLM-driven search
│ ├── vectorless.py # Mode 3: PageIndex tree structure
│ ├── graph.py # Mode 4: Knowledge graph with entity extraction
│ └── hybrid_keyword.py # Mode 5: BM25 + semantic with cross-encoder
│
├── utils/ # Utility modules
│ ├── config.py # Environment configuration loader
│ ├── metrics.py # Query metrics logging
│ └── text_utils.py # Text processing utilities
│
├── templates/ # Flask HTML templates
│ └── index.html # Main application page
│
├── static/ # Static assets
│ ├── css/
│ │ └── style.css # White & coffee styling
│ └── js/
│ └── app.js # Frontend JavaScript
│
└── uploads/ # Temporary upload directory (auto-created)
-
Select a Mode from the sidebar
- Traditional RAG: Fast vector similarity search
- Agentic RAG: Adaptive multi-step reasoning
- Vectorless RAG: Hierarchical tree navigation
- Graph RAG: Knowledge graph exploration with entity linking
- Hybrid Keyword RAG: Combined BM25 and semantic with re-ranking
-
Upload a Document (PDF, Markdown, or TXT)
- Drag & drop or click to select
- Click "Index Document" to process
-
Chat with Your Document
- Type questions in the input field
- View processing status indicators
- Expand metrics and reasoning traces for details
-
Manage Documents
- See indexed documents in the sidebar
- Clear all documents with one click
- Documents persist across sessions in ChromaDB
- Uses ChromaDB vector store with cosine similarity
- Local sentence-transformer embeddings (
all-MiniLM-L6-v2) - No API calls for embeddings - fully local
- Simple and fast retrieval
- LLM decides search strategy dynamically
- Can expand search with neighboring chunks
- Multiple reasoning steps with trace
- More adaptive but requires more tokens
- Documents converted to hierarchical tree
- LLM navigates like a human (ToC → Sections → Content)
- No embeddings or vector DB needed
- Fully explainable retrieval paths
- LLM extracts entities and relationships from documents
- Builds knowledge graph using NetworkX
- Entity-aware search with graph traversal
- Explains reasoning through entity connections
- Combines semantic embeddings with BM25 keyword search
- Uses spaCy for NLP feature extraction (NER, keywords)
- Reciprocal Rank Fusion (RRF) merges search results
- Cross-Encoder re-ranking for improved relevance
| Endpoint | Method | Description |
|---|---|---|
/ |
GET | Main application page |
/api/status |
GET | Get document status for all modes |
/api/upload |
POST | Upload and index a document |
/api/chat |
POST | Send chat query |
/api/clear |
POST | Clear indexed documents |
/api/metrics |
GET | Get query metrics history |
- Total Tokens - Prompt + Completion token count
- Retrieval Time - Time to find relevant context
- Total Latency - End-to-end response time
- Reasoning Trace - Step-by-step agent decisions (Modes 2 & 3)
FLASK_DEBUG=1 python app.pyTo clear all indexed documents:
- Click "Clear All Documents" button in sidebar, or
- Delete the
chroma_db/directory and restart
The sentence-transformer model downloads automatically on first use. If you have issues:
python -c "from sentence_transformers import SentenceTransformer; SentenceTransformer('all-MiniLM-L6-v2')"- Verify your
.envfile has correctOPENAI_API_BASEandOPENAI_API_KEY - For Ollama, ensure the server is running (
ollama serve) - Check that the model name in
MODEL_IDexists on your API
Change the port in app.py:
app.run(debug=True, host='0.0.0.0', port=5001) # Change to 5001 or any available portMIT License
- PageIndex inspiration from VectifyAI
- Sentence Transformers for local embeddings
- ChromaDB for vector storage
- Flask for web framework