Skip to content

Repository files navigation

πŸ€– AI Companion

A private, self-hosted AI assistant with persistent memory, voice interaction, and agentic capabilities β€” accessible as a PWA from any device, including instant launch via the iPhone Action Button.

Docker Python License


πŸš€ Features

Category Capabilities
πŸŽ™οΈ Voice Interface Real-time Speech-to-Text via Faster Whisper (GPU-accelerated), natural TTS responses
🧠 Persistent Memory Long-term memory with ChromaDB vector embeddings β€” your companion remembers everything
πŸ’¬ Conversational AI Powered by OpenRouter with support for multiple LLM backends (GPT-4o, Claude, Llama, etc.)
πŸ”§ Agentic Tools Code execution sandbox, web research, email integration, and extensible tool system
πŸ“± PWA Installable Progressive Web App with offline support and iPhone Action Button integration
🐳 Fully Dockerized One-command deployment with GPU passthrough β€” zero host pollution

πŸ› οΈ Tech Stack

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                  Frontend                     β”‚
β”‚  React / Vite  ->  PWA  ->  Web Audio API     β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                  Backend                      β”‚
β”‚  FastAPI  ->  Faster Whisper  ->  ChromaDB    β”‚
β”‚  SQLite   ->  OpenRouter SDK  ->  TTS Engine  β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚               Infrastructure                  β”‚
β”‚  Docker Compose  ->  NVIDIA Container Toolkit β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ“‹ Prerequisites

Requirement Version Notes
Docker Desktop 24.0+ or Docker Engine on Linux
NVIDIA Container Toolkit Latest Required for GPU-accelerated STT
NVIDIA GPU 6 GB+ VRAM For Whisper large-v3; smaller models need less
Git 2.0+ For cloning the repository

Note

GPU is required for local Whisper STT. If you don't have an NVIDIA GPU, you can switch STT_MODE to api in your .env to use a cloud-based STT provider instead.


⚑ Quick Start

# 1. Clone the repository
git clone https://github.com/your-username/AI_Companion.git
cd AI_Companion

# 2. Create your environment file
cp .env.example .env

# 3. Edit .env with your API keys
#    At minimum, set OPENROUTER_API_KEY
nano .env   # or use your preferred editor

# 4. Build and launch
docker compose up --build

# 5. Open the app
#    Frontend:  http://localhost:3000
#    Backend:   http://localhost:8000/docs  (Swagger UI)

Tip

Use docker compose up --build -d to run in detached mode. View logs with docker compose logs -f.


πŸ”Œ API Endpoints

The backend exposes a RESTful API at http://localhost:8000. Full interactive documentation is available at /docs (Swagger UI) and /redoc (ReDoc).

Method Endpoint Description
POST /api/chat Send a text message and receive an AI response
POST /api/voice/transcribe Upload audio for Speech-to-Text transcription
POST /api/voice/synthesize Convert text to speech audio
GET /api/memory/search Search long-term memory by semantic query
POST /api/memory/add Manually add an entry to long-term memory
GET /api/conversations List all conversation sessions
GET /api/conversations/{id} Retrieve a specific conversation history
DELETE /api/conversations/{id} Delete a conversation session
POST /api/tools/execute Execute code in the sandboxed environment
POST /api/tools/research Perform web research on a topic
GET /api/health Health check and system status

πŸ“² PWA Installation

Desktop (Chrome / Edge)

  1. Navigate to http://localhost:3000
  2. Click the install icon in the address bar
  3. Click Install

iOS (Safari)

  1. Open Safari and navigate to https://your-host:3000
  2. Tap Share β†’ Add to Home Screen
  3. (Optional) Set up the Action Button for instant voice access β€” see the iOS Action Button Setup Guide

Android (Chrome)

  1. Navigate to https://your-host:3000
  2. Tap the "Add to Home Screen" banner, or Menu β†’ Install App

PWA & Security Hardening

  • Auth Rate Limiting: The backend enforces sliding-window rate limiting on /api/auth/login (5 attempts/min) and /api/auth/refresh (20 attempts/min) returning HTTP 429 with Retry-After headers to protect against brute-force attacks and token replay.
  • Offline Fallback Experience: Built-in responsive offline.html fallback with live connection health diagnostics and automatic reconnect polling when internet connectivity drops.
  • Rich App Manifest: Enhanced PWA manifest with application shortcuts (Start Voice Chat, 3D Avatars), categories, and maskable icons for desktop and mobile install dialogs.

Testing & Verification

Run the test suites across backend and frontend inside the dev container:

# Backend pytest suite (auth, rate limiting, and security invariants)
cd backend && .venv/bin/python -m pytest

# Frontend test suite (API client, token refresh coalescing, and PWA manifest)
cd frontend && npm test

πŸ“ Project Structure

AI_Companion/
β”œβ”€β”€ backend/                  # FastAPI backend service
β”‚   β”œβ”€β”€ Dockerfile
β”‚   β”œβ”€β”€ requirements.txt
β”‚   β”œβ”€β”€ main.py               # Application entry point
β”‚   β”œβ”€β”€ api/                   # Route handlers
β”‚   β”‚   β”œβ”€β”€ chat.py
β”‚   β”‚   β”œβ”€β”€ voice.py
β”‚   β”‚   β”œβ”€β”€ memory.py
β”‚   β”‚   └── tools.py
β”‚   β”œβ”€β”€ core/                  # Business logic
β”‚   β”‚   β”œβ”€β”€ llm.py             # OpenRouter LLM integration
β”‚   β”‚   β”œβ”€β”€ stt.py             # Faster Whisper STT engine
β”‚   β”‚   β”œβ”€β”€ tts.py             # Text-to-Speech engine
β”‚   β”‚   β”œβ”€β”€ memory.py          # ChromaDB vector memory
β”‚   β”‚   └── sandbox.py         # Code execution sandbox
β”‚   └── models/                # Pydantic schemas
β”œβ”€β”€ frontend/                  # React PWA frontend
β”‚   β”œβ”€β”€ Dockerfile
β”‚   β”œβ”€β”€ package.json
β”‚   β”œβ”€β”€ vite.config.js
β”‚   β”œβ”€β”€ public/
β”‚   β”‚   └── manifest.json      # PWA manifest
β”‚   └── src/
β”‚       β”œβ”€β”€ App.jsx
β”‚       β”œβ”€β”€ components/        # UI components
β”‚       └── services/          # API client modules
β”œβ”€β”€ data/                      # Persistent data (git-ignored)
β”‚   β”œβ”€β”€ companion.db           # SQLite database
β”‚   └── chromadb/              # Vector embeddings
β”œβ”€β”€ sandbox/                   # Code execution workspace (git-ignored)
β”œβ”€β”€ docs/                      # Documentation
β”‚   └── ios-action-button-setup.md
β”œβ”€β”€ docker-compose.yml         # Container orchestration
β”œβ”€β”€ .env                       # Environment variables (git-ignored)
β”œβ”€β”€ .env.example               # Environment template (safe to commit)
β”œβ”€β”€ .gitignore
β”œβ”€β”€ LICENSE
└── README.md

πŸ—ΊοΈ Roadmap

Phase 1 β€” Foundation (current)

  • Docker infrastructure with GPU passthrough
  • FastAPI backend with health checks
  • Faster Whisper STT integration
  • OpenRouter LLM chat pipeline
  • SQLite conversation persistence
  • React PWA frontend with voice recording

Phase 2 β€” Intelligence

  • ChromaDB long-term vector memory
  • Semantic memory search and recall
  • Agentic tool system (code execution, web research)
  • TTS voice response pipeline
  • Conversation context management

Phase 3 β€” Polish & Expansion

  • iPhone Action Button integration & auto-record
  • Email sending capability
  • Multi-modal input (images, documents)
  • Custom personality and system prompts
  • Scheduled tasks and reminders
  • Plugin architecture for community tools

🀝 Contributing

Contributions are welcome! Please open an issue or submit a pull request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

πŸ“„ License

This project is licensed under the MIT License β€” see the LICENSE file for details.


Built with ❀️ and a healthy distrust of cloud-only AI

## Authentication setup

Before exposing this instance, follow Authentication and trusted accounts. The development ports bind to localhost. Registration and demo login are disabled by default; create the first trusted account locally, then disable registration. All operational API requests require a Bearer token. The frontend handles token attachment, refresh, and sign-out.

About

A full-stack, self-hosted AI companion platform featuring a 3D avatar with lip-sync, long-term memory, voice activity detection, and multi-agent orchestration.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages