An intelligent multi-agent system that transforms raw data into actionable insights through automated data cleaning, exploratory analysis, visualization, and comprehensive reporting.
The Intelligent Data Detective is a sophisticated multi-agent system built with LangChain and LangGraph that revolutionizes the data analysis workflow. It combines the power of Large Language Models (LLMs) with specialized agents to perform comprehensive data analysis tasks autonomously.
- 🧹 Automated Data Cleaning: Intelligently identifies and resolves data quality issues
- 📊 Exploratory Data Analysis (EDA): Performs comprehensive statistical analysis and pattern discovery
- 🎨 Dynamic Visualization: Creates relevant charts and graphs based on data insights
- 📝 Intelligent Reporting: Generates structured, narrative reports with explanations
- 🤔 Chain-of-Thought Reasoning: Provides transparent, step-by-step analytical reasoning
- 🌐 Contextual Web Search: Enriches analysis with relevant external information
Unlike traditional data analysis tools, the Intelligent Data Detective uses collaborative AI agents that work together, each specializing in different aspects of the data analysis pipeline. This approach ensures thorough, consistent, and explainable results.
flowchart LR
UI[🖥️ User Interface] --> Controller{🎯 Supervisor Agent}
Controller --> Cleaner[🧹 Data Cleaner Agent]
Controller --> Analyst[📊 Analyst Agent]
Controller --> Viz[🎨 Visualization Agent]
Controller --> Report[📝 Report Generator Agent]
Cleaner -- Cleaned Data --> Analyst
Analyst -- Insights --> Viz
Viz -- Charts --> Report
Report -- Final Report --> UI
subgraph "🔄 Workflow Orchestration"
Controller
end
subgraph "🛠️ Processing Agents"
Cleaner
Analyst
Viz
Report
end
- Role: Orchestrates the entire workflow and coordinates agent interactions
- Responsibilities:
- Manages task delegation and sequencing
- Monitors agent completion status
- Maintains global state and memory
- Makes routing decisions based on current progress
- Role: Performs intelligent data preprocessing and quality assurance
- Capabilities:
- Missing value detection and imputation strategies
- Outlier identification and handling
- Data type conversions and standardization
- Duplicate detection and removal
- Column name standardization
- Tools: pandas, numpy, custom data quality assessment tools
- Output: Cleaned dataset + detailed metadata about cleaning actions taken
- Role: Conducts comprehensive exploratory data analysis
- Capabilities:
- Descriptive statistics computation
- Correlation analysis and feature relationships
- Anomaly and pattern detection
- Hypothesis testing (normality, t-tests)
- Basic machine learning model training
- Data quality assessment
- Tools: pandas, scipy.stats, scikit-learn, custom analysis tools
- Output: Structured insights, correlations, anomalies, and visualization recommendations
- Role: Creates compelling and relevant data visualizations
- Capabilities:
- Histogram generation for distribution analysis
- Scatter plots for relationship exploration
- Correlation heatmaps for feature relationships
- Box plots for outlier visualization
- Dynamic chart selection based on data types and insights
- Tools: matplotlib, seaborn, base64 encoding for web compatibility
- Output: Generated visualizations with descriptive captions
- Role: Synthesizes all findings into comprehensive, actionable reports
- Capabilities:
- Multi-format report generation (HTML, Markdown, PDF)
- Narrative synthesis from agent findings
- Visual integration with explanatory text
- Executive summary creation
- Recommendation generation
- Tools: Jinja2 templating, xhtml2pdf, custom formatting tools
- Output: Professional reports combining text, statistics, and visualizations
- Intelligent task routing between specialized agents
- State management with checkpointing and recovery
- Memory persistence across analysis sessions
- Real-time streaming updates during processing
- Transparent step-by-step analytical reasoning
- Decision logging for reproducibility
- Explainable AI approaches to data insights
- Audit trail for all analytical decisions
- Statistical Analysis: Descriptive statistics, correlation matrices, hypothesis testing
- Quality Assessment: Missing value analysis, duplicate detection, data type validation
- Pattern Recognition: Anomaly detection, trend identification, relationship discovery
- Machine Learning: Basic classification and regression model training
- Adaptive Chart Selection: Automatically chooses appropriate visualization types
- Interactive Elements: Base64-encoded images for web integration
- Custom Styling: Configurable themes and color schemes
- Multi-format Output: PNG, SVG, and interactive plot support
- Web Search Integration: Tavily API for external context and domain knowledge
- Domain-Specific Insights: Retrieval of relevant best practices and guidelines
- Historical Learning: Memory of past analysis patterns and decisions
- Multiple Format Support: CSV, JSON, Excel file processing
- Large Dataset Optimization: Efficient memory management and processing
- Data Registry System: Centralized DataFrame management with LRU caching
- Merge and Join Operations: Multi-dataset analysis capabilities
- Advanced RAG Implementation: Vector database integration for domain knowledge
- Interactive Dashboard: Real-time data exploration interface
- Multi-dataset Comparative Analysis: Cross-dataset insights and relationships
- Advanced ML Pipeline: Automated feature engineering and model selection
- Multi-user Authentication: Secure user management and access control
- Scheduled Analysis: Automated recurring analysis jobs
- Data Privacy Controls: Field masking and anonymization capabilities
- Observability Dashboard: System monitoring and performance metrics
- Real-time Data Streaming: Live data analysis capabilities
- Predictive Analytics: Forecasting and trend prediction models
- Natural Language Queries: Conversational data exploration
- Custom Agent Development: User-defined specialized agents
- Python 3.10+: Primary programming language
- LangChain: LLM application framework and tool integration
- LangGraph: Multi-agent workflow orchestration
- OpenAI GPT-4: Primary language model for reasoning and analysis
- pandas: Data manipulation and analysis
- numpy: Numerical computing and array operations
- scipy: Statistical analysis and hypothesis testing
- scikit-learn: Machine learning algorithms and model training
- matplotlib: Comprehensive plotting library
- seaborn: Statistical data visualization
- Jinja2: Template engine for report generation
- xhtml2pdf: PDF report generation from HTML
- Tavily API: Web search and external context retrieval
- FAISS/Chroma: Vector database for RAG implementations (future)
- Memory Persistence: LangGraph's built-in state management
- File I/O: Multiple format support (CSV, JSON, Excel)
- Jupyter Notebooks: Interactive development environment
- Pydantic: Data validation and serialization
- asyncio: Asynchronous processing capabilities
- Docker: Containerization support (planned)
- Python 3.10 or higher
- OpenAI API key
- Tavily API key (optional, for web search features)
-
Clone the repository:
git clone https://github.com/dhar174/intelligent_data_detective.git cd intelligent_data_detective -
Set up environment variables:
export OPENAI_API_KEY="your-openai-api-key" export TAVILY_API_KEY="your-tavily-api-key" # Optional
-
Install dependencies:
pip install langchain langchain-core langchain-openai langchain_experimental langgraph pip install pandas numpy scipy scikit-learn matplotlib seaborn pip install pydantic python-dotenv tiktoken openpyxl xhtml2pdf pip install tavily-python chromadb joblib
The current completion baseline is the committed patched notebook:
IntelligentDataDetective_beta_v5_patched.ipynb.
The proof run IDD_run_run_default_id-20260504-1338-b3079aea completed on the deterministic retail_orders dataset with:
validate_run.pyscore 12/12validate_artifact_quality.pyscore 9/9- native structured-output markers for all major nodes
- 3/3 visualization fan-in
- canonical
final_report.html,final_report.md, andfinal_report.pdf - no recovery/final-hop/path-normalization warnings and no marker
.txtartifacts
-
Run the patched notebook through the live runner:
export OPENAI_API_KEY="your-openai-api-key" export IDD_NOTEBOOK="IntelligentDataDetective_beta_v5_patched.ipynb" export IDD_SAMPLE_DATASET="retail_orders" python run_notebook_live.py
-
Validate the latest run:
python validate_run.py --latest --log-path notebook_run_log.txt --window 180 python validate_artifact_quality.py --latest
-
Run no-key regression checks before expensive notebook proofs:
python -m pytest test_validate_run.py tests/unit tests/integration -q python -m flake8 validate_run.py validate_artifact_quality.py test_validate_run.py --max-line-length=120 --extend-ignore=E203,W503
These checks verify validator behavior, the committed W14 patched notebook markers, and the importable no-key core without requiring API credentials.
-
Regenerate the patched notebook after notebook behavior changes:
python _patch_notebook.py python -m pytest test_validate_run.py -q
Do not hand-edit IntelligentDataDetective_beta_v5_patched.ipynb; edit _patch_notebook.py and regenerate it.
You can also open the patched notebook manually:
jupyter notebook IntelligentDataDetective_beta_v5_patched.ipynb# Analyze customer sentiment and feedback patterns
analysis_prompt = """
Analyze this customer review dataset to identify:
1. Overall sentiment distribution
2. Common complaint themes
3. Product rating correlations
4. Seasonal trends in feedback
"""
# Configure analysis for notebook execution
inputs = {
"user_prompt": analysis_prompt,
"df_ids": [review_data_id],
"messages": []
}
config = {"configurable": {"thread_id": "review-analysis", "user_id": "analyst"}}
# Run analysis using the notebook's compiled graph
for result in data_detective_graph.stream(inputs, config=config, stream_mode="updates"):
print(result)# Comprehensive sales data exploration
sales_analysis = """
Examine sales performance data for:
1. Revenue trends over time
2. Top performing products/categories
3. Seasonal patterns and anomalies
4. Customer segmentation insights
"""
# Configure analysis for notebook execution
inputs = {
"user_prompt": sales_analysis,
"df_ids": [sales_data_id],
"messages": []
}
config = {"configurable": {"thread_id": "sales-analysis", "user_id": "analyst"}}
# Run analysis using the notebook's compiled graph
for result in data_detective_graph.stream(inputs, config=config, stream_mode="updates"):
print(result)# Medical data analysis with privacy considerations
healthcare_prompt = """
Analyze patient readmission data while maintaining privacy:
1. Risk factor identification
2. Readmission pattern analysis
3. Treatment effectiveness correlation
4. Demographic trend analysis
"""
# Configure analysis for notebook execution
inputs = {
"user_prompt": healthcare_prompt,
"df_ids": [patient_data_id],
"messages": []
}
config = {"configurable": {"thread_id": "healthcare-analysis", "user_id": "analyst"}}
# Run analysis using the notebook's compiled graph
for result in data_detective_graph.stream(inputs, config=config, stream_mode="updates"):
print(result)We welcome contributions from the community! Here's how you can help:
- Use the GitHub issue tracker
- Include detailed reproduction steps
- Provide sample data (anonymized) when possible
- Describe the use case and expected behavior
- Consider contributing implementation ideas
- Discuss design approaches in issues before major changes
# Clone and setup development environment
git clone https://github.com/dhar174/intelligent_data_detective.git
cd intelligent_data_detective
# Install development dependencies
pip install -e .
pip install pytest black flake8 mypy
# Run tests
pytest tests/
# Code formatting
black .
flake8 .- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
This project is licensed under the MIT License - see the LICENSE file for details.
- LangChain Team: For the incredible LLM application framework
- LangGraph Developers: For the multi-agent orchestration capabilities
- OpenAI: For providing powerful language models
- Python Data Science Community: For the amazing ecosystem of tools
- Documentation: Check out our detailed documentation
- Technical Specification: See our comprehensive tech spec
- Issues: Report bugs and request features via GitHub Issues
🔍 Transform your data into insights with the power of AI! 🔍