Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

embedding-based-similarity-engine

A lightweight, efficient retrieval engine built with ChromaDB. This project demonstrates how to ingest documents, transform them into high-dimensional vectors, and perform similarity searches using Squared Euclidean distance (L² norm).

🚀 Overview

The Embedding-Based Similarity Engine implements a clean and modular Retrieval-Augmented Generation (RAG)-style retrieval pipeline using ChromaDB as the vector store.

It allows you to:

  • Ingest raw documents
  • Convert them into embeddings
  • Persist them in a vector database
  • Run similarity searches using natural language queries
  • Inspect similarity scores via Squared Euclidean distance
  • This project is designed to be simple, extensible, and transparent - ideal for experimentation, prototyping, or educational purposes.

✨ Features

📥 Automated Ingestion

  • Loads raw documents
  • Splits them into semantically meaningful chunks
  • Stores embeddings in a persistent ChromaDB collection

🔎 Vector Search

  • Converts user queries into embeddings
  • Retrieves the most relevant document chunks

📏 Distance Metrics Transparency

  • Uses Squared Euclidean Distance (L² norm)
  • Returns raw distance scores for every retrieved document
  • Lower scores = higher similarity

🗂 Project Structure

.
├── data_ingestion.py              # Handles document loading, text chunking, and database insertion
|── embeddings_handler.py          # Manages embedding model logic to convert text into numerical vectors
├── embedding_explorer.py          # Main entry point; runnable app that accepts user queries and displays results
├── .env                           # Environment variables
├── requirements.txt
└── README.md

How It Works

The engine follows a standard RAG (Retrieval-Augmented Generation)-style retrieval pipeline:

1️⃣ Ingestion

Documents are broken into smaller chunks to:

  • Preserve semantic focus
  • Improve embedding quality
  • Increase retrieval precision

2️⃣ Embedding

Each text chunk is converted into a high-dimensional vector:

v∈Rn

This embedding captures the semantic meaning of the text.

3️⃣ Search

When a user submits a query:

  • The query is converted into an embedding vector
  • The engine compares it against stored document vectors
  • The most similar chunks are retrieved

4️⃣ Distance Calculation

Similarity is computed using Squared Euclidean Distance:

d(u, v)^2 = sum from i=1 to n of (u_i - v_i)^2

Expanded form:

d(u, v)^2 = (u_1 - v_1)^2 + (u_2 - v_2)^2 + ... + (u_n - v_n)^2

Where:

  • u = query embedding
  • v = document embedding
  • n = embedding dimensionality

📉 Interpretation

  • Lower distance → Higher similarity
  • Higher distance → Lower similarity

🛠 Installation

git clone https://github.com/kittenbytes00/embedding-based-similarity-engine.git
cd embedding-based-similarity-engine
pip install -r requirements.txt

▶️ Usage

Step 1: Ingest Data

python data_ingestion.py

Step 2: Run the Explorer

python embedding_explorer.py

📄 License

  • MIT License - free to use, modify and distribute.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages