You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A lightweight, efficient retrieval engine built with ChromaDB. This project demonstrates how to ingest documents, transform them into high-dimensional vectors, and perform similarity searches using Squared Euclidean distance (L² norm).
🚀 Overview
The Embedding-Based Similarity Engine implements a clean and modular Retrieval-Augmented Generation (RAG)-style retrieval pipeline using ChromaDB as the vector store.
It allows you to:
Ingest raw documents
Convert them into embeddings
Persist them in a vector database
Run similarity searches using natural language queries
Inspect similarity scores via Squared Euclidean distance
This project is designed to be simple, extensible, and transparent - ideal for experimentation, prototyping, or educational purposes.
✨ Features
📥 Automated Ingestion
Loads raw documents
Splits them into semantically meaningful chunks
Stores embeddings in a persistent ChromaDB collection
🔎 Vector Search
Converts user queries into embeddings
Retrieves the most relevant document chunks
📏 Distance Metrics Transparency
Uses Squared Euclidean Distance (L² norm)
Returns raw distance scores for every retrieved document
Lower scores = higher similarity
🗂 Project Structure
.
├── data_ingestion.py # Handles document loading, text chunking, and database insertion
|── embeddings_handler.py # Manages embedding model logic to convert text into numerical vectors
├── embedding_explorer.py # Main entry point; runnable app that accepts user queries and displays results
├── .env # Environment variables
├── requirements.txt
└── README.md
How It Works
The engine follows a standard RAG (Retrieval-Augmented Generation)-style retrieval pipeline:
1️⃣ Ingestion
Documents are broken into smaller chunks to:
Preserve semantic focus
Improve embedding quality
Increase retrieval precision
2️⃣ Embedding
Each text chunk is converted into a high-dimensional vector:
v∈Rn
This embedding captures the semantic meaning of the text.
3️⃣ Search
When a user submits a query:
The query is converted into an embedding vector
The engine compares it against stored document vectors
The most similar chunks are retrieved
4️⃣ Distance Calculation
Similarity is computed using Squared Euclidean Distance: