hub / github.com/PromtEngineer/localGPT

github.com/PromtEngineer/localGPT @main sqlite

540 symbols 1,695 edges 88 files 244 documented · 45%

README

LocalGPT - Private Document Intelligence Platform

<a href="https://x.com/engineerrprompt">
  <img src="https://img.shields.io/badge/Follow%20on%20X-000000?style=for-the-badge&logo=x&logoColor=white" alt="Follow on X" />
</a>
<a href="https://discord.gg/tUDWAFGc">
  <img src="https://img.shields.io/badge/Join%20our%20Discord-5865F2?style=for-the-badge&logo=discord&logoColor=white" alt="Join our Discord" />
</a>

🚀 What is LocalGPT?

LocalGPT is a fully private, on-premise Document Intelligence platform. Ask questions, summarise, and uncover insights from your files with state-of-the-art AI—no data ever leaves your machine.

More than a traditional RAG (Retrieval-Augmented Generation) tool, LocalGPT features a hybrid search engine that blends semantic similarity, keyword matching, and Late Chunking for long-context precision. A smart router automatically selects between RAG and direct LLM answering for every query, while contextual enrichment and sentence-level Context Pruning surface only the most relevant content. An independent verification pass adds an extra layer of accuracy.

The architecture is modular and lightweight—enable only the components you need. With a pure-Python core and minimal dependencies, LocalGPT is simple to deploy, run, and maintain on any infrastructure.The system has minimal dependencies on frameworks and libraries, making it easy to deploy and maintain. The RAG system is pure python and does not require any additional dependencies.

▶️ Video

Watch this video to get started with LocalGPT.

Home	Create Index	Chat

✨ Features

Utmost Privacy: Your data remains on your computer, ensuring 100% security.
Versatile Model Support: Seamlessly integrate a variety of open-source models via Ollama.
Diverse Embeddings: Choose from a range of open-source embeddings.
Reuse Your LLM: Once downloaded, reuse your LLM without the need for repeated downloads.
Chat History: Remembers your previous conversations (in a session).
API: LocalGPT has an API that you can use for building RAG Applications.
GPU, CPU, HPU & MPS Support: Supports multiple platforms out of the box, Chat with your data using CUDA, CPU, HPU (Intel® Gaudi®) or MPS and more!

📖 Document Processing

Multi-format Support: PDF, DOCX, TXT, Markdown, and more (Currently only PDF is supported)
Contextual Enrichment: Enhanced document understanding with AI-generated context, inspired by Contextual Retrieval
Batch Processing: Handle multiple documents simultaneously

🤖 AI-Powered Chat

Natural Language Queries: Ask questions in plain English
Source Attribution: Every answer includes document references
Smart Routing: Automatically chooses between RAG and direct LLM responses
Query Decomposition: Breaks complex queries into sub-questions for better answers
Semantic Caching: TTL-based caching with similarity matching for faster responses
Session-Aware History: Maintains conversation context across interactions
Answer Verification: Independent verification pass for accuracy
Multiple AI Models: Ollama for inference, HuggingFace for embeddings and reranking

🛠️ Developer-Friendly

RESTful APIs: Complete API access for integration
Real-time Progress: Live updates during document processing
Flexible Configuration: Customize models, chunk sizes, and search parameters
Extensible Architecture: Plugin system for custom components

🎨 Modern Interface

Intuitive Web UI: Clean, responsive design
Session Management: Organize conversations by topic
Index Management: Easy document collection management
Real-time Chat: Streaming responses for immediate feedback

🚀 Quick Start

Note: The installation is currently only tested on macOS.

Prerequisites

Python 3.8 or higher (tested with Python 3.11.5)
Node.js 16+ and npm (tested with Node.js 23.10.0, npm 10.9.2)
Docker (optional, for containerized deployment)
8GB+ RAM (16GB+ recommended)
Ollama (required for both deployment approaches)

NOTE

Before this brach is moved to the main branch, please clone this branch for instalation:

git clone -b localgpt-v2 https://github.com/PromtEngineer/localGPT.git
cd localGPT

Option 1: Docker Deployment

# Clone the repository
git clone https://github.com/PromtEngineer/localGPT.git
cd localGPT

# Install Ollama locally (required even for Docker)
curl -fsSL https://ollama.ai/install.sh | sh
ollama pull qwen3:0.6b
ollama pull qwen3:8b

# Start Ollama
ollama serve

# Start with Docker (in a new terminal)
./start-docker.sh

# Access the application
open http://localhost:3000

Docker Management Commands:

# Check container status
docker compose ps

# View logs
docker compose logs -f

# Stop containers
./start-docker.sh stop

Option 2: Direct Development (Recommended for Development)

# Clone the repository
git clone https://github.com/PromtEngineer/localGPT.git
cd localGPT

# Install Python dependencies
pip install -r requirements.txt

# Key dependencies installed:
# - torch==2.4.1, transformers==4.51.0 (AI models)
# - lancedb (vector database)
# - rank_bm25, fuzzywuzzy (search algorithms)
# - sentence_transformers, rerankers (embedding/reranking)
# - docling (document processing)
# - colpali-engine (multimodal processing - support coming soon)

# Install Node.js dependencies
npm install

# Install and start Ollama
curl -fsSL https://ollama.ai/install.sh | sh
ollama pull qwen3:0.6b
ollama pull qwen3:8b
ollama serve

# Start the system (in a new terminal)
python run_system.py

# Access the application
open http://localhost:3000

System Management:

# Check system health (comprehensive diagnostics)
python system_health_check.py

# Check service status and health
python run_system.py --health

# Start in production mode
python run_system.py --mode prod

# Skip frontend (backend + RAG API only)
python run_system.py --no-frontend

# View aggregated logs
python run_system.py --logs-only

# Stop all services
python run_system.py --stop
# Or press Ctrl+C in the terminal running python run_system.py

Service Architecture: The run_system.py launcher manages four key services: - Ollama Server (port 11434): AI model serving - RAG API Server (port 8001): Document processing and retrieval - Backend Server (port 8000): Session management and API endpoints - Frontend Server (port 3000): React/Next.js web interface

Option 3: Manual Component Startup

# Terminal 1: Start Ollama
ollama serve

# Terminal 2: Start RAG API
python -m rag_system.api_server

# Terminal 3: Start Backend
cd backend && python server.py

# Terminal 4: Start Frontend
npm run dev

# Access at http://localhost:3000

Detailed Installation

1. Install System Dependencies

Ubuntu/Debian:

sudo apt update
sudo apt install python3.8 python3-pip nodejs npm docker.io docker-compose

macOS:

brew install python@3.8 node npm docker docker-compose

Windows:

# Install Python 3.8+, Node.js, and Docker Desktop
# Then use PowerShell or WSL2

2. Install AI Models

Install Ollama (Recommended):

# Install Ollama
curl -fsSL https://ollama.ai/install.sh | sh

# Pull recommended models
ollama pull qwen3:0.6b          # Fast generation model
ollama pull qwen3:8b            # High-quality generation model

3. Configure Environment

# Copy environment template
cp .env.example .env

# Edit configuration
nano .env

Key Configuration Options:

# AI Models (referenced in rag_system/main.py)
OLLAMA_HOST=http://localhost:11434

# Database Paths (used by backend and RAG system)
DATABASE_PATH=./backend/chat_data.db
VECTOR_DB_PATH=./lancedb

# Server Settings (used by run_system.py)
BACKEND_PORT=8000
FRONTEND_PORT=3000
RAG_API_PORT=8001

# Optional: Override default models
GENERATION_MODEL=qwen3:8b
ENRICHMENT_MODEL=qwen3:0.6b
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
RERANKER_MODEL=answerdotai/answerai-colbert-small-v1

4. Initialize the System

# Run system health check
python system_health_check.py

# Initialize databases
python -c "from backend.database import ChatDatabase; ChatDatabase().init_database()"

# Test installation
python -c "from rag_system.main import get_agent; print('✅ Installation successful!')"

# Validate complete setup
python run_system.py --health

🎯 Getting Started

1. Create Your First Index

An index is a collection of processed documents that you can chat with.

Using the Web Interface:

Open http://localhost:3000
Click "Create New Index"
Upload your documents (PDF, DOCX, TXT)
Configure processing options
Click "Build Index"

Using Scripts:

# Simple script approach
./simple_create_index.sh "My Documents" "path/to/document.pdf"

# Interactive script
python create_index_script.py

Using API:

# Create index
curl -X POST http://localhost:8000/indexes \
  -H "Content-Type: application/json" \
  -d '{"name": "My Index", "description": "My documents"}'

# Upload documents
curl -X POST http://localhost:8000/indexes/INDEX_ID/upload \
  -F "files=@document.pdf"

# Build index
curl -X POST http://localhost:8000/indexes/INDEX_ID/build

2. Start Chatting

Once your index is built:

Create a Chat Session: Click "New Chat" or use an existing session
Select Your Index: Choose which document collection to query
Ask Questions: Type natural language questions about your documents
Get Answers: Receive AI-generated responses with source citations

3. Advanced Features

Custom Model Configuration

# Use different models for different tasks
curl -X POST http://localhost:8000/sessions \
  -H "Content-Type: application/json" \
  -d '{
    "title": "High Quality Session",
    "model": "qwen3:8b",
    "embedding_model": "Qwen/Qwen3-Embedding-4B"
  }'

Batch Document Processing

# Process multiple documents at once
python demo_batch_indexing.py --config batch_indexing_config.json

API Integration

import requests

# Chat with your documents via API
response = requests.post('http://localhost:8000/chat', json={
    'query': 'What are the key findings in the research papers?',
    'session_id': 'your-session-id',
    'search_type': 'hybrid',
    'retrieval_k': 20
})

print(response.json()['response'])

🔧 Configuration

Model Configuration

LocalGPT supports multiple AI model providers with centralized configuration:

Ollama Models (Local Inference)

OLLAMA_CONFIG = {
    "host": "http://localhost:11434",
    "generation_model": "qwen3:8b",        # Main text generation
    "enrichment_model": "qwen3:0.6b"       # Lightweight routing/enrichment
}

External Models (HuggingFace Direct)

EXTERNAL_MODELS = {
    "embedding_model": "Qwen/Qwen3-Embedding-0.6B",           # 1024 dimensions
    "reranker_model": "answerdotai/answerai-colbert-small-v1", # ColBERT reranker
    "fallback_reranker": "BAAI/bge-reranker-base"             # Backup reranker
}

Pipeline Configuration

LocalGPT offers two main pipeline configurations:

Default Pipeline (Production-Ready)

"default": {
    "description": "Production-ready pipeline with hybrid search, AI reranking, and verification",
    "storage": {
        "lancedb_uri": "./lancedb",
        "text_table_name": "text_pages_v3",
        "bm25_path": "./index_store/bm25"
    },
    "retrieval": {
        "retriever": "multivector",
        "search_type": "hybrid",
        "late_chunking": {"enabled": True},
        "dense": {"enabled": True, "weight": 0.7},
        "bm25": {"enabled": True}
    },
    "reranker": {
        "enabled": True,
        "type": "ai",
        "strategy": "rerankers-lib",
        "model_name": "answerdotai/answerai-colbert-small-v1",
        "top_k": 10
    },
    "query_decomposition": {"enabled": True, "max_sub_queries": 3},
    "verification": {"enabled": True},
    "retrieval_k": 20,
    "contextual_enricher": {"enabled": True, "window_size": 1}
}

Fast Pipeline (Speed-Optimized)

Extension points exported contracts — how you extend this code

Props (Interface)

(no doc)

src/components/SessionIndexInfo.tsx

Step (Interface)

(no doc)

src/lib/api.ts

MarkdownProps (Interface)

(no doc)

src/components/Markdown.tsx

ChatMessage (Interface)

src/components/IndexPicker.tsx

ChatSession (Interface)

src/components/IndexWizard.tsx

ChatRequest (Interface)

(no doc)

src/lib/api.ts

Core symbols most depended-on inside this repo

system_health_check.py

run

called by 17

rag_system/agent/loop.py

get_user_input

called by 16

create_index_script.py

send_json_response

called by 15

rag_system/api_server.py

send_json_response

called by 14

rag_system/api_server_with_progress.py

update

called by 12

rag_system/utils/batch_processor.py

Shape

Method 284

Function 168

Class 50

Interface 38

Languages

Python67%

TypeScript33%

Modules by API surface

backend/server.py38 symbols

src/lib/api.ts37 symbols

backend/database.py25 symbols

run_system.py22 symbols

rag_system/api_server_with_progress.py22 symbols

rag_system/agent/loop.py18 symbols

rag_system/pipelines/retrieval_pipeline.py17 symbols

rag_system/utils/batch_processor.py16 symbols

src/components/ui/dropdown-menu.tsx15 symbols

rag_system/indexing/representations.py14 symbols

rag_system/api_server.py14 symbols

demo_batch_indexing.py12 symbols

Dependencies from manifests, versioned

@eslint/eslintrc3 · 1×

@radix-ui/react-avatar1.1.10 · 1×

@radix-ui/react-dropdown-menu2.1.15 · 1×

@radix-ui/react-scroll-area1.2.9 · 1×

@radix-ui/react-separator1.1.7 · 1×

@radix-ui/react-slot1.2.3 · 1×

@tailwindcss/postcss4 · 1×

@types/node20 · 1×

@types/react19 · 1×

@types/react-dom19 · 1×

class-variance-authority0.7.1 · 1×

clsx2.1.1 · 1×

For agents

$ claude mcp add localGPT \
  -- python -m otcore.mcp_server <graph>

⬇ download graph artifact