An enterprise-ready, stateful conversational AI assistant powered by LangGraph, LangChain, and Streamlit. Designed with a modular ReAct agent architecture, persistent SQLite checkpointing, multi-thread conversation management, external tool execution (real-time web search and mathematics), and optimized local Hugging Face inference with 4-bit quantization.
- Key Features
- System Architecture
- Project Directory Structure
- Prerequisites
- Installation & Quickstart
- Configuration Guide
- Database Schema & Memory Management
- Troubleshooting & FAQ
- Contributing
- License
- Stateful Multi-Turn Conversations: Powered by LangGraph's compiled
StateGraphandSqliteSavercheckpointer. State is persisted across server reboots, page refreshes, and thread switches. - Dynamic Tool Calling & Agentic Loop: Features automated ReAct conditional loops (
tools_condition) that trigger specialized capabilities when required:- Live Web Search: Real-time internet search via DuckDuckGo.
- Arithmetic Calculator: Precise numeric computations with error boundary handling.
- Extensible Registry: Automatic discovery of any tool files dropped into
backend/tools/.
- Hardware-Optimized Local Inference:
- Configured for
Qwen/Qwen2.5-3B-Instruct, balancing quality, reasoning, and resource efficiency. - 4-Bit NF4 Quantization via
bitsandbytesreducing VRAM requirements to ~2GB. - Seamless toggle between local Hugging Face pipelines (
HuggingFacePipeline) and cloud-hosted inference endpoints (HuggingFaceEndpoint).
- Configured for
- Interactive Multi-Thread UI:
- Sidebar conversation history with instant creation, renaming, selection, and deletion of threads.
- Interactive "New Chat" and "Rename" modal dialogs with auto-focus form controls.
- Live temperature slider for fine-tuning model creativity vs. tool determinism.
- Real-Time Token Streaming:
- Streaming responses rendered token-by-token using
chatbot.stream(..., stream_mode='messages'). - Visual indicators for active tool calls (
🛠️ Using tool...) and execution results (✔️ Tool returned...).
- Streaming responses rendered token-by-token using
- Production-Grade Fault Tolerance:
- Centralized Precision Logging: Every log message shows the exact file, function, and line number.
- SQLite WAL (Write-Ahead Logging) mode and exponential backoff retries to prevent database deadlocks.
- Lazy factory instantiation for the chatbot model to recover gracefully from misconfigurations.
flowchart TD
subgraph UI ["Frontend Layer (Streamlit)"]
A[User Input] --> B[Streamlit Session State]
B --> C[Sidebar: Thread Manager & Settings]
B --> D[Chat Message Renderer & Token Streamer]
end
subgraph LangGraph ["Agent Engine (LangGraph)"]
E[START] --> F[chat_node]
F --> G{tools_condition}
G -- Tool Requested --> H[ToolNode: Execution]
H --> F
G -- Final Answer --> I[END]
end
subgraph Memory ["Persistence Layer (SQLite)"]
J[(chatbot.db)]
K[Thread Metadata: threads]
L[Graph Checkpoints: SqliteSaver]
J --- K
J --- L
end
subgraph Inference ["Inference & Tools"]
M[Hugging Face Model\nLocal / API]
N[DuckDuckGo Search]
O[Arithmetic Calculator]
end
D <==>|Invoke / Stream| LangGraph
LangGraph <==>|State Checkpoints| Memory
F <==>|Generate Tokens / Tool Calls| M
H <==>|Execute Tool| N
H <==>|Execute Tool| O
The project has been modularized into highly focused sub-packages:
ChatBot-in-LangGraph/
│
├── backend/ # Core agent logic and inference services
│ ├── bot/ # ChatBot class and lazy cache factory
│ ├── config/ # Config loading, validation, and CLI selector
│ ├── db/ # SQLite connection and thread operations
│ ├── graph/ # LangGraph state machine, nodes, and parser
│ ├── model/ # Hugging Face local & API model loader
│ ├── tools/ # Auto-discovering tool registry
│ │ ├── __init__.py # Discovers tools using pkgutil
│ │ ├── calculator.py # Arithmetic calculator tool
│ │ └── search.py # DuckDuckGo search tool
│ ├── constants.py # Single source of truth for magic strings
│ ├── logger.py # Centralized precise logging setup
│ └── models.json # Catalog of AI models, specs, and requirements
│
├── frontend/ # Presentation and user interface
│ ├── state/ # Session state, conversation loaders, thread management
│ └── ui/ # Sidebar, dialogs, and chat stream rendering
│
├── .gitignore # Git tracking exemptions
├── app.py # Main Streamlit application entry point
├── CodeStructure.md # Project coding conventions and style guide
├── config.json # General project configurations & settings
├── README.md # Comprehensive project documentation
├── requirements.txt # Categorized project dependencies (CUDA 12.6)
└── run.bat # Automated Windows launcher and venv manager
Before starting, ensure your system meets the following specifications:
| Requirement | Minimum | Recommended |
|---|---|---|
| Operating System | Windows 10/11, Ubuntu 20.04+, or macOS | Windows 11 / Linux (64-bit) |
| Python Version | Python 3.10 | Python 3.11 or 3.12 |
| RAM | 8 GB System Memory | 16 GB System Memory |
| GPU (Optional) | CPU-only mode supported | NVIDIA GPU with 4GB+ VRAM (CUDA 12.0+) |
The repository provides a self-healing run.bat script that verifies your Python installation, provisions a virtual environment, installs dependencies, and launches the application:
- Clone the repository:
git clone https://github.com/Mr-Rup/ChatBot-in-LangGraph.git cd ChatBot-in-LangGraph - Double-click
run.bat(or execute it via Command Prompt / PowerShell):.\run.bat
- The script will automatically configure the
.myenvenvironment and open the Streamlit interface athttp://localhost:8501.
-
Clone the Repository:
git clone https://github.com/Mr-Rup/ChatBot-in-LangGraph.git cd ChatBot-in-LangGraph -
Create and Activate a Virtual Environment:
- Windows:
python -m venv .myenv .myenv\Scripts\activate
- Linux / macOS:
python3 -m venv .myenv source .myenv/bin/activate
- Windows:
-
Install PyTorch with CUDA Support: If you have an NVIDIA GPU, install PyTorch with CUDA 12.6 acceleration first:
pip install torch torchvision --extra-index-url https://download.pytorch.org/whl/cu126
-
Install Remaining Dependencies:
pip install -r requirements.txt
-
Configure Environment Variables (Optional but Recommended): Create a
.envfile in the root directory and add your API keys (see the Configuration Guide below for exactly what keys to add and where to get them). -
Launch the Application:
streamlit run app.py
The project adopts a strict separation between sensitive credentials, application runtime settings, and AI model specifications.
The project uses a .env file in the root directory for all sensitive API keys. You must create this file manually. None of the keys are strictly mandatory, but they unlock advanced API models and premium search capabilities.
# ==========================================
# Language Models
# ==========================================
# Gemini API Key (Required for gemini-3.6-flash and gemini-3.6-pro)
# Get it here: https://aistudio.google.com/app/apikey
GEMINI_API_KEY=your_gemini_api_key_here
# Hugging Face Access Token (Required for gated repos like Llama-3.2)
# Get it here: https://huggingface.co/settings/tokens
HUGGINGFACEHUB_API_TOKEN=your_huggingface_api_token_here
# ==========================================
# Smart Search Fallback Tool
# ==========================================
# Tavily AI Search (Highest quality AI search results)
# Get a free key here: https://tavily.com
TAVILY_API_KEY=your_tavily_api_key_here
# Google Custom Search (High quality web search fallback)
# Get API Key: https://console.cloud.google.com/apis/credentials
GOOGLE_API_KEY=your_google_api_key_here
# Get CSE ID: https://programmablesearchengine.google.com/
GOOGLE_CSE_ID=your_google_cse_id_here
# ==========================================
# Observability
# ==========================================
# LangSmith API Key (Required only if langsmith tracing is enabled in config.json)
# Get it here: https://smith.langchain.com
LANGCHAIN_API_KEY=your_langsmith_api_key_hereAll non-sensitive application settings (cache directories, database paths, LangSmith toggles, and system prompts) are controlled centrally in config.json:
{
"active_model": "qwen-2.5-3b",
"llm_cache_dir": null,
"database_path": "chatbot.db",
"langsmith": {
"tracing": false,
"project": "ChatBot-LangGraph",
"endpoint": "https://api.smith.langchain.com"
},
"system_prompt": "You are a highly capable AI assistant with access to external tools. You MUST use these tools when asked to perform math, search, or look up information. Do NOT refuse to use tools. Do NOT perform calculations yourself. Always output the correct JSON format to invoke the tool when needed."
}All AI model definitions, hardware specifications, and parameters live cleanly in backend/models.json:
{
"qwen-2.5-3b": {
"name": "Qwen 2.5 3B Instruct",
"repo_id": "Qwen/Qwen2.5-3B-Instruct",
"model_type": "local",
"task": "text-generation",
"temperature": 0.1,
"max_new_tokens": 512,
"specs": {
"parameters": "3.09B",
"vram_required": "~2.0 GB (4-bit quantized)",
"ram_required": "8 GB",
"tool_support": "High (Native function calling & tool precision)"
},
"description": "Recommended. Outstanding balance of deep reasoning, concise answers, and tool-use reliability."
}
}- Qwen 2.5 3B Instruct (
qwen-2.5-3b): Default. Outstanding reasoning and reliable tool-calling precision (~2GB VRAM). - TinyLlama 1.1B Chat (
tiny-llama-1.1b): Ultra-lightweight and fast, runs on virtually any PC/laptop (~1GB VRAM). - Qwen 2.5 1.5B Instruct (
qwen-2.5-1.5b): Great balance of speed and instruction following (~1.2GB VRAM). - Qwen 2.5 0.5B Instruct (
qwen-2.5-0.5b): Ultra-compact, ideal for CPU-only and testing (<1GB VRAM). - Llama 3.2 1B Instruct (
llama-3.2-1b): Edge-optimized multilingual model (Requires HF token in.env). - Phi 3.5 Mini Instruct (
phi-3.5-mini): Microsoft's strong mathematical and reasoning model (~2.5GB VRAM).
- At Launch (Interactive CLI): When executing
run.bat, an interactive menu prompts you to either keep the active model or select a new one from the list. - In Streamlit UI: The sidebar automatically displays the active model's name, parameter count, hardware requirements, and tool capability.
- Manually: Change the
"active_model"key directly inconfig.json.
The project uses a self-registering tool discovery system. Adding new capabilities to the chatbot is fully automated.
- Create a new file in the
backend/tools/directory (e.g.,backend/tools/stock_price.py). - Define your tool using the
@tooldecorator. - Expose a
TOOLSlist at the bottom of the file.
from langchain_core.tools import tool, BaseTool
@tool
def get_stock_price(ticker: str) -> dict:
"""Fetch the latest stock price for a given stock ticker symbol."""
return {"ticker": ticker, "price": 182.50}
# The tools/__init__.py auto-discovers this list!
TOOLS: list[BaseTool] = [get_stock_price]That's it! The system will automatically find your tool, bind it to the LLM, and execute it when requested. You don't need to change any other file.
All session and conversation history is persisted in a local SQLite database (chatbot.db) configured with WAL (Write-Ahead Logging) mode and automated backoff retries for concurrent read/write stability.
-
threads(Managed bybackend/db/threads.py):thread_id(TEXT, PRIMARY KEY): Unique thread identifier (e.g.,thread1,thread2).thread_name(TEXT): User-defined label displayed in the sidebar.updated_at(TIMESTAMP): Tracks the most recently active conversations.
-
checkpoints&writes(Managed bySqliteSaver):- Stores serialized LangGraph states, checkpoint snapshots, and message histories.
- Deleting a thread via the UI triggers a cascading cleanup that removes both thread metadata and associated checkpoints to save disk space.
1. CUDA Out of Memory (OOM) error during local model loading
- Solution: Verify
bitsandbytesis installed to ensure 4-bit quantization is active. If your GPU has less than 4GB VRAM, switch'model_type': 'api'inbackend/models.jsonor use a smaller base model likeQwen/Qwen2.5-1.5B-InstructorQwen/Qwen2.5-0.5B-Instruct.
2. DuckDuckGo Search fails with rate limit errors
- Solution: Ensure
duckduckgo-searchis up to date:pip install --upgrade duckduckgo-search
3. SQLite database is locked (OperationalError)
- Solution: WAL mode is enabled and there is automatic backoff-retry logic in place. However, ensure no external SQLite browser has locked the file in exclusive mode.
4. Changing model download cache directory
- Solution: Edit
config.jsonand change"llm_cache_dir": nullto the absolute path of your choice (e.g.,"S:/ollama_models"). If leftnull, Hugging Face will use its platform default (~/.cache/huggingface).
Contributions, issues, and feature requests are welcome!
- Fork the Project
- Create your Feature Branch (
git checkout -b feature/AmazingFeature) - Commit your Changes (
git commit -m 'Add some AmazingFeature') - Push to the Branch (
git push origin feature/AmazingFeature) - Open a Pull Request
Please adhere to the coding standards described in CodeStructure.md.
Distributed under the MIT License. See LICENSE for more information.