Find precise answers from your files and repos using local embeddings, smart query planning, and embedded vector search.
RAGnarΕk helps developers, knowledge workers, and enterprise teams search, summarize, review, and answer questions over local documents, repositories, and the active VS Code workspace β with privacy and compliance in mind. Use it fully offline with local Transformers.js embeddings and LanceDB storage, or enable optional LLM-based planning and evaluation via VS Code Copilot models without any external API key for advanced query decomposition and result assessment.
Why install?
- Fast, private semantic search over PDFs, Markdown, HTML, and code
- Enterprise-friendly: per-topic stores, file-based persistence, and secure token handling for private repos
- Agentic query planning and evaluation: optionally use LLMs for decomposition, iterative refinement, and answer evaluation
- Include workspace context: surface relevant open files, symbols, and code snippets to enrich answers
- Code-review assistance: apply retrieved guidelines and documentation to review your code and get actionable suggestions
- Embedded LanceDB vector store β no external servers required
- Works offline with local Transformers.js embedding models
- Run embeddings locally: Use Transformers.js models (ONNX/wasm) without external APIs.
- Local model picker: Load models from
ragnarok.localModelPathand switch models in the tree view. - Offline & private: Keep embeddings and inference on-device for privacy and compliance.
- Default model included: Ships with
Xenova/all-MiniLM-L6-v2by default for fast, 384-dimension embeddings.
RAGnarΕk supports multiple embedding providers via a pluggable backend system:
| Mode | Setting value | Description |
|---|---|---|
| Auto | auto (default) |
Tries registered backends in order; uses first available |
| VS Code LM | vscodeLM |
Uses the proposed vscode.lm.computeEmbeddings API (requires a registered provider such as GitHub Copilot) |
| HuggingFace | huggingface |
Local Transformers.js ONNX/WASM inference β fully offline, no external services |
| Remote | remote |
OpenAI or Ollama-compatible embedding API (MCP server only) |
Configuration:
ragnarok.embeddingBackendβ selectauto,vscodeLM,huggingface, or any registered backend nameragnarok.embeddingVscodeModelIdβ (optional) specific VS Code LM model ID; leave blank to auto-select
Prerequisites for VS Code LM embeddings:
- VS Code Insiders (or any build that supports the proposed embeddings API)
"enabledApiProposals": ["embeddings"]in the extension manifest (already configured)- An embeddings provider registered at runtime (e.g., GitHub Copilot with embeddings support)
β οΈ Known limitation: Thevscode.lm.computeEmbeddingsAPI is a proposed API and may not be available on stable VS Code builds. When usingautomode, the extension shows user-visible warning/info notifications and falls back to HuggingFace if the API is unavailable.
To use the VS Code Language Model embeddings API (vscode.lm.computeEmbeddings) you must enable proposed APIs for this extension and start the VS Code with the --enable-proposed-api flag referencing the extension id.
You can also add the same flag as a runtime argument in your VS Code.
Ctrl+Shift+P -> "Preferences: Configure Runtime Arguments" and add the following to the argv.json:
{
"enable-proposed-api": ["hyorman.ragnarok"]
}Notes:
- If you run VS Code remotely (WSL/Containers), run the
code/code-insiderscommand on the host where the Extension Host will run. - After enabling proposed APIs restart the Extension Development Host.
- A proposed API requires a runtime provider (e.g., GitHub Copilot) β ensure the provider is installed and active.
- Intelligent Query Decomposition: Automatically breaks complex queries into sub-queries
-- LLM-Powered Planning: Uses Copilot (VS Code LM API) models such as
gpt-4ofor advanced reasoning (Copilot required; no external API key). LLM usage is optional - Heuristic Fallback: Works without LLM using rule-based planning
- Iterative Refinement: Confidence-based iteration for high-quality results
- Parallel/Sequential Execution: Smart execution strategy based on query complexity
- Hybrid Search (recommended): Combines vector + keyword (90%/10% weights, configurable)
- Vector Search: Pure semantic similarity using embeddings
- BM25 Search: Pure keyword search using Okapi BM25 algorithm (no embeddings needed)
- Cross-Encoder Reranking: Optional second-stage reranking over any strategy's candidates
- Position Boosting: Keywords near document start weighted higher
- Result Explanations: Human-readable scoring breakdown for all strategies
- Multi-Format Support: PDF, Markdown, HTML, plain text, GitHub repositories
- Semantic Chunking: Automatic strategy selection (markdown/code/recursive)
- Structure Preservation: Maintains heading hierarchy and context
- Batch Processing: Multi-file upload with progress tracking
- GitHub Integration: Load entire repositories from GitHub.com or GitHub Enterprise Server
- LangChain Loaders: Industry-standard document loading
- LanceDB: Embedded vector database with file-based persistence (no server needed)
- Cross-Platform: Works on Windows, macOS, Linux, and ARM
- Per-Topic Stores: Efficient isolation and management
- Serverless: Truly embedded, like SQLite for vectors
- Caching: Optimized loading and reuse
- Configuration View: See agentic settings at a glance
- Embedding Model Picker: Tree view lists curated + local models (from
ragnarok.localModelPath) with download status; click to switch - Statistics Display: Documents, chunks, store type, model info
- Progress Tracking: Real-time updates during processing
- Rich Icons: Visual hierarchy with emojis and theme icons
- Comprehensive Logging: Debug output at every step
- Type-Safe: Full TypeScript with strict mode
- Error Handling: Robust error recovery throughout
- Async-Safe: Mutex locks prevent race conditions
- Configurable: 15+ settings for customization
- Shared Core, Separate Data: VS Code and MCP delegate memory operations to the same core
MemoryService, but use separate storage roots. VS Code uses its extensionglobalStorageUri; MCP usesRAGNAROK_STORAGE_DIR. There is no cross-host data sharing or automatic migration. - Native VS Code Tools: the extension contributes exactly three language-model tools β
ragQueryto search a topic,ragTopicto list topics or inspect one topic's statistics and documents, andragMemoryfor scoped memory operations.ragQueryandragTopicare read-only.ragMemorystores, recalls, and forgets individual memories, but it has no reset action: wiping memory outright is Reset Memory in the RAG sidebar's Memory section, behind a modal confirmation, just as creating, renaming, exporting, importing, and deleting topics are sidebar actions. The three tools' input schemas are generated from the canonical JSON Schema contracts in@ragnarok/corebynpm run tools:manifestand drift-checked bynpm run tools:manifest:check. - Persistent Project Memory: Store and recall facts, preferences, conventions, and context across sessions β scoped to workspace or git branch
- Automatic Git Branch Detection: Memories can be scoped per branch via
GitBranchDetector, auto-detecting the current branch from the working directory - Vector-Based Recall + Entity Graph: Memories are embedded and stored in a dedicated LanceDB instance; an entity graph (graphology) tracks relationships between extracted concepts
- LLM-Powered Entity Extraction: Optionally extracts entities (facts, preferences, concepts, tools, conventions) from stored memories; the graph stays empty when no LLM provider is configured
- Markdown Export: Automatically generates a
memories.mdfile summarizing stored memories for human review - MCP Integration: Exposed as the
rag_memorytool with store, recall, forget, stats, list, decay, history, promote, link, and community operations.communitiesclusters the memory entity graph and requires an LLM provider β without one the tool says so instead of returning an empty result. Memory TTL is supported; reservedauto:memories are hidden unless explicitly requested. Memory is written and recalled only through explicitrag_memorycalls β there is no automatic query-time recall or write-back.
Graphs exist only in the memory subsystem. There is no document knowledge
graph, no entity extraction over ingested documents, and no graph or
graph_hybrid retrieval strategy.
The rag_memory_visualize tool exports the local user's own memory graph. It
accepts exactly
{ source: "memory", memoryScope: "workspace", maxNodes? } or
{ source: "memory", memoryScope: "branch", branch, maxNodes? } and returns the
deterministic ragnarok.graph.visualization.v1 document. The default is 500
nodes, the accepted range is 1 through 2,000, and output is capped at 10,000
edges and the MCP response-byte limit. Failures surface as
GRAPH_VISUALIZATION_RECORD_TOO_LARGE or GRAPH_VISUALIZATION_FAILED; an empty
or unknown scope returns an empty document rather than fabricated data.
Documents include full persisted node/edge descriptions, provenance,
confidence, scope/branch fields, and arbitrary metadata, but never embedding
vectors. In VS Code, run RAGnarok: Show Memory Graph to choose workspace or
current-branch memory and open the interactive command webview. MCP Apps hosts
instead load the self-contained ui://ragnarok/graph resource as
text/html;profile=mcp-app via modern _meta.ui.resourceUri. Both surfaces use
the shared renderer and deterministic graph document, but each reads its own
host's storage. There is no cross-host data sharing. See the
MCP server graph contract.
git clone https://github.com/hyorman/ragnarok.git
cd ragnarok
npm install
npm run compile
# Press F5 to run in development modecode --install-extension ragnarok-0.1.6.vsix- Default:
Xenova/all-MiniLM-L6-v2 - Offline/local: set
ragnarok.localModelPathto a folder containing Transformers.js-compatible models (each model in its own subfolder). The tree view will list those models alongside curated ones; click any entry to load it. - When you change the embedding model, existing topics keep their original embeddingsβcreate a new topic if you need to ingest with the new model.
Cmd/Ctrl+Shift+P β RAG: Create New Topic
Enter name (e.g., "React Docs") and optional description.
Cmd/Ctrl+Shift+P β RAG: Add Document to Topic
Select topic, then choose one or more files. The extension will:
- Load documents using LangChain loaders
- Apply semantic chunking
- Generate embeddings
- Store in vector database
Supported formats: .pdf, .md, .html, .txt
Cmd/Ctrl+Shift+P β RAG: Add GitHub Repository to Topic
Or right-click a topic in the tree view and select the GitHub icon. You can:
- GitHub.com or GitHub Enterprise Server: Choose between public GitHub or your organization's GitHub Enterprise Server
- Enter repository URL:
- GitHub.com:
https://github.com/facebook/react - GitHub Enterprise:
https://github.company.com/team/project
- GitHub.com:
- Specify branch (defaults to
main) - Configure ignore patterns (e.g.,
*.test.js, docs/*) - Add access token for private repositories (see Token Management below)
The extension will recursively load all files from the repository and process them just like local documents.
Note: Supports GitHub.com and GitHub Enterprise Server only. The repository must be accessible from your network. For other Git hosting services (GitLab, Bitbucket, etc.), clone the repository locally and add it as local files.
For accessing private repositories, RAGnarΕk securely stores GitHub access tokens per host using VS Code's Secret Storage API.
Add a Token:
Cmd/Ctrl+Shift+P β RAG: Add GitHub Token
- Enter the GitHub host (e.g.,
github.com,github.company.com) - Paste your GitHub Personal Access Token (PAT)
- The token is securely stored and automatically used for that host
List Saved Tokens:
Cmd/Ctrl+Shift+P β RAG: List GitHub Tokens
Shows all hosts with saved tokens (tokens themselves are never displayed).
Remove a Token:
Cmd/Ctrl+Shift+P β RAG: Remove GitHub Token
Select a host to remove its stored token.
Export a Topic:
Cmd/Ctrl+Shift+P β RAG: Export Topic
Or select a topic in the tree view and select the export icon. This creates a portable archive containing:
- Topic metadata (name, description)
- Vector embeddings and documents
- Model configuration
Exported topics can be shared with teammates or imported into other workspaces.
Import a Topic:
Cmd/Ctrl+Shift+P β RAG: Import Topic
Or click the import icon in the tree view title bar. Select an exported topic archive to restore it into your workspace.
Rename a Topic:
Cmd/Ctrl+Shift+P β RAG: Rename Topic
Or select a topic in the tree view and click the edit icon.
How to Create a GitHub PAT:
- Go to GitHub Settings β Developer settings β Personal access tokens β Tokens (classic)
- Click "Generate new token (classic)"
- Select the
reposcope - Generate and copy the token
- Use the "RAG: Add GitHub Token" command to save it
Benefits:
- β Tokens stored securely in VS Code's Secret Storage (not in settings.json)
- β Support for multiple GitHub hosts (GitHub.com + multiple Enterprise servers)
- β Automatic token selection based on repository URL
- β No need to enter token every time you add a repository
RAGnarΕk supports read-only access to shared team knowledge bases via the ragnarok.commonDatabasePath setting.
Setup:
- Export topics from a source workspace
- Place exported topic archives in a shared location (network drive, shared folder)
- Configure
ragnarok.commonDatabasePathto point to this folder:
{
"ragnarok.commonDatabasePath": "/path/to/shared/rag-databases"
}Benefits:
- β Share curated knowledge bases across teams
- β Read-only topics prevent accidental modification
- β Centralized documentation and policy storage
- β Works with any file-sharing system
Note: Topics from common database path appear in the tree view but cannot be deleted or modified.
Open Copilot Chat (@workspace)
Type: @workspace #ragQuery What is [your question]?
The RAG tool will:
- Match your topic semantically
- Decompose complex queries (if agentic mode enabled)
- Perform hybrid retrieval
- Return ranked results with context
Clear Model Cache:
Cmd/Ctrl+Shift+P β RAG: Clear Model Cache
Removes cached embedding models. Useful when switching models or troubleshooting.
Clear Database:
Cmd/Ctrl+Shift+P β RAG: Clear Database
Refresh Topics:
Cmd/Ctrl+Shift+P β RAG: Refresh Topics
Reloads the topic tree view. Useful after importing topics or external changes.
{
// Path to local Transformers.js embedding model folder
"ragnarok.localModelPath": "",
// Number of results to return
"ragnarok.topK": 5,
// Chunk size for splitting documents
"ragnarok.chunkSize": 512,
// Chunk overlap for context preservation
"ragnarok.chunkOverlap": 50,
// Retrieval strategy: vector, hybrid, bm25
"ragnarok.retrievalStrategy": "hybrid",
// Path to shared/common RAG database (read-only topics)
"ragnarok.commonDatabasePath": ""
}Note: GitHub access tokens are now managed via secure Secret Storage, not settings.json. See GitHub Token Management section.
{
// Maximum refinement iterations (1-10)
"ragnarok.maxIterations": 3,
// Confidence threshold (0-1) for stopping iteration
"ragnarok.confidenceThreshold": 0.7,
// LLM model: gpt-4o, gpt-4o-mini, gpt-3.5-turbo
"ragnarok.llmModel": "gpt-4o",
// Include workspace context (selected code, active file, imports, symbols)
"ragnarok.includeWorkspaceContext": true
}Set ragnarok.localModelPath to point at a folder that already contains compatible Transformers.js models (one subfolder per modelβe.g., an ONNX export downloaded ahead of time). Entries found here appear in the tree view and can be selected directly, and this local path takes precedence over ragnarok.embeddingModel.
Available Embedding Models to Download (local, no API needed):
Xenova/all-MiniLM-L6-v2(default) - Fast, 384 dimensionsXenova/all-MiniLM-L12-v2- More accurate, 384 dimensionsXenova/paraphrase-MiniLM-L6-v2- Optimized for paraphrasingXenova/multi-qa-MiniLM-L6-cos-v1- Optimized for Q&A
The extension ships with Xenova/all-MiniLM-L6-v2 by default; to use other local models, set ragnarok.localModelPath or click the model name in tree view.
Any models you place under ragnarok.localModelPath show up in the tree view alongside these curated options (with download indicators) and can be loaded with one click.
LLM Models (when agentic planning is enabled): models are available via VS Code Copilot / LM API (no external API key required).
gpt-4o(default) - Most intelligentgpt-4o-mini- Faster, still capablegpt-3.5-turbo- Fastest, most economical
RAGnarΕk is organized as an npm workspaces monorepo with four packages:
copilot-rag/
βββ packages/
β βββ core/ # @ragnarok/core β portable RAG engine (no VS Code dependency)
β βββ graph-ui/ # @ragnarok/graph-ui β private shared graph renderer/build
β βββ vscode/ # @ragnarok/vscode β VS Code extension adapters and UI
β βββ mcp-server/ # @ragnarok/mcp-server β MCP server for CLI/TUI/GUI agents
βββ test/ # VS Code extension test infrastructure and fixtures
βββ assets/ # Extension icon and bundled embedding models
βββ scripts/ # Build and packaging helpers
| Package | Description |
|---|---|
@ragnarok/core |
Loaders, chunkers, embeddings, retrievers, agents, stores β all platform-agnostic with dependency injection |
@ragnarok/graph-ui |
Private browser source and build that generates the VS Code JS/CSS and self-contained MCP App HTML |
@ragnarok/vscode |
VS Code adapters (IConfigProvider, ILogger, INotifier, ILLMProvider), commands, tree view, and extension entry point |
@ragnarok/mcp-server |
Exposes RAG and memory tools via the Model Context Protocol β works with any MCP-compatible agent (stdio transport only) |
RAGnarok 0.4.0 serves MCP protocol 2026-07-28 only. Clients must use
server/discover or modern version negotiation; legacy initialize is
rejected. There is no compatibility mode and no Mcp-Session-Id.
Stdio is the only transport. The HTTP transport, shared deployment mode,
bearer roles, and upload/download handles were removed; the server is a child
process of one MCP client, running as the user who spawned it. Environment
variables belonging to the removed transport are rejected at startup, with an
error naming every one that was set. Cacheable discovery,
list, and resource-read results advertise ttlMs=0 and cacheScope=private.
See the MCP server guide for the complete tool
surface and configuration.
MCP server settings live in config.json in the storage directory, which the
server generates on first run. That file is the only place they are set β a key
present in it pins your value, a key absent uses the current built-in default,
and there is no environment variable for any of them. The environment carries
only credentials, the two bootstrap paths, and the two one-shot switches. The
key table lists both sets.
(These are the MCP server's settings. The VS Code extension is configured
separately through the ragnarok.* settings above.)
npm install # Install all workspace dependencies
npm run compile # Build all packages (tsc -b)
npm run test:all # Run all tests (core β vscode β mcp-server)
npm run test:core # Run core package tests only
npm test # Run VS Code extension tests only
npm run test:mcp # Run MCP server tests only
npm run bench:smoke # Fast deterministic retrieval/reranker gate
npm run bench:release # Pinned release benchmark; missing inputs fail
npm run test:docs # Validate canonical documentation links/contracts
npm run lint # Lint all packages
npm run format # Format all source and test files
npm run clean # Clean all build artifactsThe MCP server exposes these tools to any MCP-compatible agent:
| Tool | Description |
|---|---|
rag_query |
Query a topic with agentic RAG (supports all retrieval strategies); shares one executor with the VS Code ragQuery tool |
rag_ingest |
Add content to a topic from local files, one public HTTP(S) page, or an allowlisted GitHub/GHES repository |
rag_topic |
Manage topics: list, stats (statistics plus indexed documents), create, rename, export, and import (list and stats share one implementation with the VS Code ragTopic tool) |
rag_delete_topic |
Delete a topic after explicit confirmation |
rag_remove_document |
Remove a document and reconcile its chunks |
rag_memory |
Store, recall, forget, list, or get stats for project memories (workspace/branch-scoped); shares its input normalizer and MemoryService with the VS Code ragMemory tool |
rag_reset_memory |
Reset incompatible or unwanted standalone memory after confirmation. It has no VS Code counterpart: a headless agent has no sidebar to click, so this stays a tool guarded by confirm: true |
rag_memory_visualize |
Return a deterministic memory graph document and associate the MCP App |
That is the complete surface: 8 tools, all registered unconditionally on every
connection. There are no roles and no capability tiers β the client already runs
with the owner's authority. For the MCP server, the embedding model, the
reranker, and the LLM provider are configured exclusively through config.json
and expose no tools. Parameters and error
contracts are in the MCP server guide.
Version 0.4.0 uses storage format v2 and .rag archive format 2.0. New empty installations initialize automatically. Non-empty 0.3/unversioned storage fails closed and must be converted with the supported offline migrator; VS Code offers a preview before migration and never silently resets it. See MIGRATION.md. Embedding fingerprints are persisted per topic and memory store so incompatible semantic spaces are rejected even when dimensions happen to match.
Single-writer constraint: Only one process (VS Code window, MCP server instance, or CLI tool) may access a storage directory at a time. A second process fails fast instead of silently corrupting data. See the architecture for the complete concurrency and locking model.
- Architecture
- Storage migration
- Operations and recovery
- Security
- Benchmark gates
- Release evidence and publication
Release evidence is truthful by construction: required jobs are recorded as
passed, failed, or unrun. Docker runtime and all six installed VSIX platform
combinations are release blockers until their designated CI environments
execute them; a local compile or package build does not imply those gates
passed. The release benchmark also exits nonzero with status: "blocked" when
child-process peak RSS, isolated index time, or exact package-size
measurements are absent; deterministic smoke tests do not stand in for those
declared measurements.
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β VS Code Extension β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β βββββββββββββββ ββββββββββββββββ ββββββββββββββ β
β β Commands β β Tree View β β RAG Tool β β
β β (UI) β β (UI) β β (Copilot) β β
β βββββββ¬ββββββββ ββββββββ¬ββββββββ βββββββ¬βββββββ β
β β β β β
β βββββββ΄ββββββββββββββββββ΄ββββββββββββββββββ΄βββββββ β
β β Topic Manager β β
β β (Topic lifecycle, caching, coordination) β β
β βββββββ¬βββββββββββββββββββββββββββββββββββ¬ββββββββ β
β β β β
β βββββββ΄ββββββββββ ββββββββ΄ββββββββ β
β β Document β β RAG Agent β β
β β Pipeline β β (Orchestr.) β β
β ββ¬ββββββββββ¬βββββ ββ¬ββββββββββ¬ββββ β
β β β β β β
β βββ΄βββββ ββββ΄βββββ ββββββββ΄βββ ββββββ΄ββββ β
β βLoaderβ βChunkerβ β Planner β βRetriev.β β
β β β β β β β β β β
β ββββ¬ββββ βββββ¬ββββ ββββββ¬βββββ βββββ¬βββββ β
β β β β β β
β βββ΄ββββββββββ΄βββββ ββββββ΄βββββββββββ΄βββββ β
β β Embedding β β Vector Store β β
β β Service β β (LanceDB) β β
β β (Local Models) β β (Embedded DB) β β
β ββββββββββββββββββ ββββββββββββββββββββββ β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
ββββββββ΄ββββββββ
β LangChain.js β
β (Foundation) β
ββββββββββββββββ
User Query: "Compare React hooks vs class components"
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 1. Topic Matching (Semantic Similarity) β
β β Finds best matching topic β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 2. Query Planning (LLM or Heuristic) β
β Complexity: complex β
β Sub-queries: β
β - "React hooks features and usage" β
β - "React class components features" β
β Strategy: parallel β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 3. Hybrid Retrieval (for each sub-query) β
β Vector search: 70% weight β
β Keyword search: 30% weight β
β β Returns ranked results β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 4. Iterative Refinement (if enabled) β
β Check confidence: 0.65 < 0.7 β
β β Refine query and retrieve again β
β Check confidence: 0.78 β₯ 0.7 β β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 5. Result Processing β
β - Deduplicate by content hash β
β - Rank by score β
β - Limit to topK β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
Return: Ranked results with metadata
User uploads: document1.pdf, document2.md
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 1. Document Loading (LangChain Loaders) β
β PDF: PDFLoader β
β MD: TextLoader β
β HTML: CheerioWebBaseLoader β
β β Returns Document[] with metadata β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 2. Semantic Chunking β
β Strategy selection: β
β - Markdown: MarkdownTextSplitter β
β - Code: RecursiveCharacterTextSplitter β
β - Other: RecursiveCharacterTextSplitter β
β β Preserves headings and structure β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 3. Embedding Generation (Batched) β
β Model: Xenova/all-MiniLM-L6-v2 (local) β
β Batch size: 32 chunks β
β β Generates 384-dim vectors β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
βββββ΄βββββββββββββββββββββββββββββββββββββββββ
β 4. Vector Storage β
β LanceDB embedded database β
β β Stores embeddings + metadata β
βββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
Complete: Documents ready for retrieval
Performance depends on CPU architecture, model revision, corpus, storage, and Node version. Reproducible smoke and release-grade benchmark commands, pinned inputs, quality thresholds, latency/memory/package budgets, and the reviewed baseline update process are documented in docs/BENCHMARKS.md. Historical approximate timings live in docs/BENCHMARK-HISTORY.md and are not treated as release evidence.
- Use local embeddings for privacy and no API costs
- Enable agent caching (automatic per topic)
- Adjust chunk size based on document type
- Use simple mode for fast queries
- Batch document uploads for efficiency
- Measure your corpus β capacity and latency are bounded by local storage, memory, native dependencies, and workload shape
| Problem | Solution |
|---|---|
| "No embeddings provider registered" | Ensure a provider (e.g., GitHub Copilot) is installed and active. Set ragnarok.embeddingBackend to huggingface as a workaround. |
| "Proposed API not enabled" | The vscode.lm.computeEmbeddings API requires "enabledApiProposals": ["embeddings"] in the extension manifest. Use VS Code Insiders for full support. |
| VS Code LM embedding dimension mismatch | Switching backends may change the embedding dimension. Existing vector stores need re-indexing after backend changes. Delete the topic and re-add documents. |
| Fallback warnings appearing frequently | If you see repeated "falling back to HuggingFace" messages, either set ragnarok.embeddingBackend to huggingface explicitly, or check that your VS Code LM provider is running. |
| Model not found in VS Code LM | Verify the model ID in ragnarok.embeddingVscodeModelId matches one listed in vscode.lm.embeddingModels. Leave blank to auto-select. |
npm testOpen an issue before large changes and include the relevant compile, lint, test, benchmark, migration, or packaging evidence with the pull request.
git clone https://github.com/hyorman/ragnarok.git
cd ragnarok
npm install
npm run watch # Watch mode for developmentMIT License - see LICENSE for details
Built with:
- LangChain.js - Document processing framework
- Transformers.js - Local embeddings
- LanceDB - Embedded vector database
- VS Code Extension API - Extension platform
- VS Code LM API - Copilot integration
Made with β€οΈ by the hyorman
β Star us on GitHub if you find this useful!