An autonomous, configuration-free data quality and self-healing data pipeline powered by local AI Agents and an extensible RAG (Retrieval-Augmented Generation) infrastructure.
Enterprise data pipelines frequently break down due to silent data anomalies—such as transactional calculation mismatches, broken ledger records, or context-specific data violations. Traditional validation tools rely on hardcoded validation scripts or rigid database assertions. These systems are fragile and shatter the moment a database schema is mutated or expanded.
AegisData addresses this problem by replacing hardcoded validation models with a dynamic AI Agent that scans database schemas on its own, groups rows logically based on transactional markers, matches business rules using a semantic RAG vector network, and exports isolated, human-reviewable .sql repair patches.
- Stage 1: Ingestion & Inference — The user drops a target CSV file into the workspace. The engine reads raw text strings to infer optimal SQL data types and construct autogenerated DDL parameters.
- Stage 2: Sandbox Initialisation — The application drops any pre-existing schemas and compiles a localized sandbox database inside an isolated SQLite instance.
- Stage 3: Dynamic Reflection — A decoupled database inspector connects to the target metadata catalog and reflects active table parameters on the fly without relying on static codebase structures.
- Stage 4: Adaptive Window Chunking — The orchestrator scans column names for session tracking attributes. If tracking flags are detected, it bundles correlated row clusters into transactional blocks instead of standard row-by-row streams.
- Stage 5: RAG Context Enrichment — The system searches a persistent ChromaDB instance for validation rule matches that align semantically with the table name and row payload.
- Stage 6: Autonomous Evaluation — The system packages the schema, matched templates, and row structures into a rigid context layout. A local Ollama instance processes the request using deterministic JSON enforcement constraints.
- Stage 7: Isolated Export — Flagged calculation anomalies and target SQL adjustments are compiled and exported as a clean text patch script for explicit user validation.
- Zero-Config Schema Reflection: Uses SQLAlchemy's inspection layer to dynamically read database configurations, constraints, and schemas at runtime. The application requires zero manual code configuration changes if a database column changes.
- Adaptive Window Chunking: Programmatically reads and targets relational anomalies by evaluating table layouts. If session keywords (like transaction_id or session_id) are present, it groups row clusters into unified transactional payloads to maintain contextual meaning.
- IaC Validation Framework (RAG Engine): Validation rules are managed as discrete, human-readable items inside an external JSON tracking configuration. A localized ChromaDB layer synchronizes and searches these rules using vector embeddings.
- 100% Free, Private Local Inference: Uses the OpenAI SDK mapped directly to a local Ollama server running Llama 3.2 / 3. No corporate data leaves your local machine, and there are zero cloud token costs.
- Deterministic Validation Outputs: Employs explicit system prompts paired with a 0.0 temperature setting to enforce strict, schema-compliant JSON payloads.
- Explicit Human Approval (HITL): Discovered anomalies and suggested updates are aggregated and exported into a separate .sql patch script for user validation, avoiding risky automated executions.
config/
└── validation_rules.json # Configuration file for tracking validation rules
demo/
├── sample_data.csv # Seed data file containing intentional ledger anomalies
└── run_demo.py # CSV Ingestion, data type parsing, and sandbox builder
src/
├── __init__.py
├── main.py # Main orchestrator pipeline script
├── extractor/
│ ├── chunker.py # Heuristic engine deciding single-row vs. transactional sessions
│ └── schema.py # Dynamic SQLAlchemy database structure reflection
├── rag/
│ └── vector_store.py # Local ChromaDB configurations and vector upsert operations
├── agent/
│ ├── llm_reasoner.py # OpenAI SDK execution mapping into Ollama API
│ └── prompts.py # System prompts and JSON-enforcement layouts
└── utils/
└── logger.py # Logging configuration utilities
.env # Environment configuration secrets management
.gitignore # System exclusions file
requirements.txt # Pinned project dependencies
Ensure you have Python 3.10+ installed. Clone the repository and install the minimal system dependencies:
git clone https://github.com/AtulSingh-Emyre/AegisData.git
cd AegisData
pip install -r requirements.txt
Create a .env file in the project root:
TARGET_TABLE_NAME=customer_transactions
1. Download and run Ollama from ollama.com.
2. Pull down your preferred lightweight model via your system terminal:
ollama run llama3.2
Execute the two-step execution script from your project root folder:
Step A: Process the CSV file, infer datatypes, and seed the Sandbox DB
python demo/run_demo.py
Step B: Run the main autonomous validation pipeline framework
python -m src.main
Upon pipeline execution, AegisData parses your transactional records, flags anomalies, and creates a local .sql patch script:
-- ==================================================
-- AegisData Autogenerated Data Optimization Patches
-- Target Table: customer_transactions
-- ==================================================
-- Anomaly: Sum of individual line_amount records for tx_1001 ($250) does not match transaction_total_header ($300).
UPDATE customer_transactions
SET transaction_total_header = 250.0
WHERE transaction_id = 'tx_1001';
- Language: Python 3.12
- Database Management: SQLAlchemy Core (Dynamic Schema Reflection)
- Vector Database: ChromaDB (Local Persistent Embedded Storage)
- AI Engine Interface: OpenAI Python Client SDK (Mapped to localhost:11434)
- Local Inference Model: Meta Llama 3.2 (3B) / Llama 3 (8B) via Ollama