Skip to content

Repository files navigation

SecondBrain AgenticRAG

A modernized Second Brain system that keeps conversational memory, routes queries through an agentic workflow, and stores searchable knowledge across text and images.

Getting Started

Prerequisites

  • Docker & Docker Compose
  • API keys for: OpenAI, Grok (optional)
  • Notion token (optional, for Notion push feature)

Setup

  1. Clone the repository:

    git clone https://github.com/yourusername/SecondBrian_AgenticRAG.git
    cd SecondBrian_AgenticRAG
  2. Create .env file with your credentials (see .env.example).

  3. Start all services:

    docker compose up -d
  4. Access the UI at http://localhost:8000

Reset & Debug Mode

Use the provided reset script to perform a clean deployment. Note: This script automatically enables DEBUG logging for all containers.

./reset_and_run.sh

Highlights

  • Agentic RAG — LangGraph-powered workflow that classifies, routes, retrieves, and refines
  • Hierarchical Memory (L1/L2/L3) — Working turns, episodic session headlines/summaries, and a durable semantic profile, all in PostgreSQL
  • Session Consolidation — Conversations are automatically summarized into headlines every few turns and at session end, so past sessions stay searchable
  • Grounded Answers — Final synthesis treats Retrieved Context as the sole source of truth about past conversations; model opinions are advisory
  • Multi-LLM Parallel Calls — Queries OpenAI and Grok simultaneously with identical context
  • Notion Push — Export conversation summaries directly to your Notion database (/tools/notion_push)

Architecture

System Architecture

Services Overview

Service Port Override (env var) Description
Frontend 8000 FRONTEND_PORT Web UI for chat interface — open http://localhost:8000
API 8001 API_PORT FastAPI + LangGraph agentic pipeline (health: /health)
Vector 8002 VECTOR_PORT Embedding generation and semantic search (health: /collection/stats)
Qdrant 6333 QDRANT_HTTP_PORT Vector database HTTP API
PostgreSQL 5432 POSTGRES_PORT Sessions, messages, and hierarchical memory

All ports are published by docker-compose.yml; set the env vars (e.g. in .env) to remap them.

Accessing from another device: the browser derives the API URL from the page's host (http://<host>:8001), so opening http://<server-ip>:8000 from a LAN device works out of the box. API_BASE_URL in .env only needs changing if the API is exposed on a different host/port than <frontend-host>:8001; a localhost value is ignored by non-localhost browsers.


Agentic Workflow

Workflow Diagram

The agentic RAG pipeline follows these stages (nodes in create_agentic_rag_graph()):

  1. Classify Session (classify_session) — LLM categorizes the session as "RAG" (History Search) or "NO_RAG" (General). NO_RAG skips straight to the model fan-out.
  2. Route (route_query) — LLM decides whether retrieval is needed (only for RAG mode).
  3. Retrieve (retrieve_context) — Calls the hierarchical memory facade (MemoryManager.retrieve), gathering profile facts (L3), headline shortlist (L2), and recent working turns (L1). No LLM used.
  4. Refine (refine_query) — If retrieval came back empty and iterations remain, an LLM rewrites the query and retrieval is retried.
  5. Select & Unfold (select_and_unfold) — Expands shortlisted session headlines into full summary bodies before synthesis.
  6. Call LLMs (call_llms) — Parallel calls to OpenAI/Grok with the injected context.
  7. Answer (generate_answer) — Synthesizes the final response, grounded in the Retrieved Context.

Agent Flow

Agent Flow

Memory Sources

Retrieval goes through a single facade — MemoryManager.retrieve (services/api-service/db/memory/manager.py) — which returns three tiers:

Tier Source Storage Scope Purpose
L1 Working conversation_messages PostgreSQL Current session Last 10 turns, oldest-first, for immediate context
L2 Episodic conversation_sessions headlines + summaries PostgreSQL Cross-session Full-text shortlist of past-session headlines (current session excluded)
L3 Semantic user_memory profile facts PostgreSQL Durable, per-user Identity/preference facts, salience-ranked under a token budget

Hierarchical Memory

How the tiers are populated and searched (see docs/hierarchical-memory-plan/ for the full design):

  • Consolidation lifecycle — After each assistant turn, /ask fires a debounced, fire-and-forget consolidation (ask_flow.py): consolidate_session no-ops until the session's message counter reaches CONSOLIDATE_EVERY (4, a constant in db/memory/constants.py), then writes a headline + summary using MEMORY_SUMMARY_MODEL (default gpt-4o-mini), with a naive text fallback so L2 never stays empty. POST /sessions/new force-consolidates the previous session before deactivating it — end-of-session is always a consolidation point.
  • Backfill — Consolidate every legacy session whose headline is missing:
    docker compose run --rm --no-deps api python -m tools.backfill_consolidate
  • Three-tier headline shortlist (L2) — shortlist_headlines (db/memory/retrieval.py) fills up to 8 slots: (1) strict full-text AND match; (2) the same terms OR-relaxed so filler words don't cause misses; (3) raw conversation_messages.content match returning the session's headline — bridging paraphrase gaps (a "cities" question can find a session whose summary never says "cities").
  • Debugging — With LOG_LEVEL=DEBUG (e.g. via ./reset_and_run.sh), every tier access logs as [L1 working] / [L2 episodic] / [L3 semantic]: docker compose logs -f api | grep '\[L'.
  • Known gap — promote_to_profile (L3 fact extraction) has no production trigger yet; wiring it into consolidation is a follow-up.

LLM Providers

All LLMs receive the full retrieved context injected into their prompts:

  • OpenAI — gpt-4o for agent reasoning (classify/route/refine/synthesize); OPENAI_MODEL for parallel calls (config default gpt-5-mini; note the graph's env fallback is gpt-4o-mini when the variable is unset)
  • Grok (GROK_MODEL, default grok-4.3) — Alternative perspective

/analyze-image uses GPT (OpenAI vision) by default; the Claude/Gemini/Grok vision paths only run when explicitly requested via selected_models. None of them are wired into the text-chat model fan-out.


AgenticRAG Mechanism

This project implements an Agentic RAG (Retrieval-Augmented Generation) system powered by LangGraph, featuring intelligent routing, hierarchical memory retrieval, parallel LLM calls, and grounded synthesis.

How It Works

The agentic workflow operates as a state machine with the following stages:

1. Classify Session (classify_session_node)

  • Purpose: Determines the intent of a new session based on the first query.
  • LLM: GPT-4o
  • Modes:
    • RAG: History Search / Context Dependent (triggers full pipeline)
    • NO_RAG: New Knowledge / General Assistance (bypasses retrieval for speed)
  • Persistence: Remembers classification for the duration of the session.

2. Route Query (route_query_node)

  • Purpose: Intelligently decides whether the query requires context retrieval
  • LLM: GPT-4o at low temperature (for consistent reasoning)
  • Logic:
    • Analyzes the user's query to determine if it needs historical context
    • Automatically forces retrieval for multi-turn conversations to maintain context
    • Outputs: should_retrieve decision flag

3. Retrieve Context (retrieve_context_node)

  • Purpose: Gathers relevant information via the hierarchical memory facade
  • No LLM used — Pure data retrieval for speed and cost efficiency
  • Implementation: acquires a DB connection and calls MemoryManager.retrieve(conn, query, user_id, session_id), which returns:
    1. Profile facts (L3, user_memory) — durable cross-session identity/preferences
    2. Headlines (L2, conversation_sessions) — shortlisted candidate summaries from other sessions
    3. Working turns (L1, conversation_messages) — recent turns from the current session
  • Each item is tagged with its source (profile / headlines / working) in retrieved_context.

4. Refine Query (refine_query_node) — Optional

  • Purpose: Improves the query if initial retrieval returns insufficient context
  • Trigger: Activated when retrieved_context is empty but iterations remain (AGENTIC_MAX_ITERATIONS)
  • LLM: GPT-4o refines the query phrasing
  • Iteration: Retries retrieval with the refined query

5. Select & Unfold (select_and_unfold_node)

  • Purpose: Expands shortlisted L2 headlines into full summary bodies before synthesis
  • Behavior: Collects headlines items from retrieved_context, unfolds their bodies, and appends them as unfolded context items
  • Crash-safe: If no headlines were retrieved (or unfolding fails), it is a pass-through that leaves retrieved_context untouched

6. Call LLMs (call_llms_node)

  • Purpose: Parallel fan-out to multiple LLM providers with injected context
  • Models: OpenAI (OPENAI_MODEL), Grok (GROK_MODEL)
  • Context Injection:
    • Retrieved context formatted into a structured prompt
    • Session history included for conversational coherence
    • Each model receives identical context for consistent comparison
  • Execution: All models called simultaneously via asyncio.gather(); per-model answers and errors tracked separately (llm_answers / llm_errors)

7. Generate Answer (generate_answer_node)

  • Purpose: Synthesizes a final, grounded answer
  • LLM: GPT-4o as the synthesis engine
  • Grounding rules (agentic/prompts.py):
    • The Retrieved Context is the only source of truth about past conversations — if model opinions conflict with it, or claim past discussions that do not appear in it, the context wins
    • If a topic is absent from the Retrieved Context, the answer states plainly that there is no record of it
  • Input: Retrieved (and unfolded) context, individual answers from all LLMs, and conversation history

State Management

The workflow uses LangGraph's StateGraph with typed state (ConversationState, agentic/state.py):

{
  "session_id": str,           # Current conversation session
  "user_id": str,              # User identifier
  "session_mode": str,         # "RAG" or "NO_RAG" classification
  "current_query": str,        # User's question (can be refined)
  "messages": List[dict],      # Conversation history (reducer-annotated: operator.add)
  "should_retrieve": bool,     # Routing decision
  "should_refine": bool,       # Whether to refine the query
  "retrieved_context": List,   # Context items tagged by memory tier/source
  "selected_models": List,     # Which LLMs to call
  "llm_answers": Dict,         # Individual model responses
  "llm_errors": Dict,          # Error tracking per model (not fed to synthesis)
  "final_answer": str,         # Synthesized response
  "response_mode": str,        # "fast" or "llm-wiki"
  "skill_prompt": str,         # Active skill instructions (llm-wiki mode)
  "agent_thoughts": List,      # Debugging/transparency log
  "iteration_count": int,      # Current iteration
  "max_iterations": int,       # Iteration limit (AGENTIC_MAX_ITERATIONS, default 2)
  "flow_timing": Dict          # Per-node / per-model timing metadata
}

Key Features

Feature Description
Session Classification Distinguishes between historical search vs general requests on first query
Conditional Routing Not all queries need retrieval—simple questions go straight to the model fan-out
Hierarchical Memory (L1/L2/L3) Working turns, episodic session headlines/summaries, and semantic profile facts through one facade
Session Consolidation Debounced headline+summary generation per turn, forced at session end, backfillable for legacy sessions
Grounded Synthesis Retrieved Context is the sole source of truth about past conversations; model opinions are advisory
Parallel LLM Calls Calls OpenAI and Grok simultaneously for diverse perspectives
Transparent Logging agent_thoughts and flow_timing track decisions and per-node latency for debugging

Flow Control

The workflow uses conditional edges for dynamic routing:

Agentic Flow

Edges: classify_session -> call_llms (NO_RAG) | route_query; route_query -> retrieve_context | call_llms | END; retrieve_context -> refine_query | select_and_unfold; refine_query -> retrieve_context; select_and_unfold -> call_llms; call_llms -> generate_answer -> END.

Why Agentic?

Traditional RAG systems follow a fixed pipeline: retrieve → generate. This agentic approach adds:

  1. Intelligence: Decides when retrieval is necessary
  2. Adaptability: Refines queries if initial retrieval fails
  3. Redundancy: Calls multiple LLMs for robustness
  4. Context-Awareness: Full conversation memory across sessions
  5. Observability: Explicit state tracking and decision logging

This architecture ensures that answers are not just generated from context, but intelligently orchestrated through a reasoning process that adapts to each query's needs.


API Endpoints

All routes live in services/api-service/main.py unless noted.

Health & Sessions

Method Endpoint Description
GET /health Service health check
POST /sessions/create Create a new conversation session
GET /sessions List user sessions
GET /sessions/{session_id} Get session detail (incl. messages)
DELETE /sessions/{session_id} Delete a session
PUT /sessions/{session_id}/title Rename a session
POST /sessions/new Start a new session (consolidates + deactivates the previous one)
GET /sessions/{session_id}/wiki-export.zip Export session as wiki bundle (zip)
GET /sessions/{session_id}/wiki-export.md Export session as a single markdown file

Chat

Method Endpoint Description
POST /ask Main question endpoint (triggers agentic RAG)
POST /analyze-image Vision analysis (GPT by default; Claude/Gemini/Grok via explicit selected_models)
POST /reset Reset a session's working memory

Knowledge & Status

Method Endpoint Description
POST /query Simple query endpoint (stub implementation)
POST /store-knowledge Store a Q&A pair in the vector database
GET /search/{query} Search stored knowledge (configurable threshold)
GET /status Status of all microservices

Vector Operations (proxy to vector-service)

Method Endpoint Description
POST /vector/search Semantic search
POST /vector/topics Topics extracted from actual Q&A sessions
POST /vector/stats Vector collection stats
GET /vector/knowledge-graph User's knowledge graph
POST /vector/by-topic User's entries filtered by topic

Tool Routers

Method Endpoint Description
POST /tools/daily_summarizer Generate daily summary (tools/summarizer.py)
POST /tools/notion_push Push a summary to Notion (tools/notion_push.py)
POST /tools/daily_summarize_and_push Summarize + Notion push combo (tools/daily_summarize_and_push.py)
— /vocabulary/* Vocabulary memory CRUD, search, spaced-repetition review, and stats (13 routes, tools/vocabulary_routes.py)

Legacy Tools (unwired)

agentic/tools.py still defines semantic_search_tool, session_history_tool, user_memories_tool, and knowledge_graph_tool, but the graph no longer uses them — retrieval goes through MemoryManager. They are kept for reference and slated for consolidation.


Environment Variables

Variable Default Description
OPENAI_API_KEY — OpenAI API key (required)
LOG_LEVEL INFO Global log level (DEBUG enables [L1]/[L2]/[L3] memory tracing)
OPENAI_MODEL gpt-5-mini OpenAI model for parallel LLM calls (config default; graph falls back to gpt-4o-mini if unset)
GEMINI_MODEL gemini-2.0-flash-exp Gemini model (summarization path + opt-in vision analysis; not in the ask fan-out)
GROK_MODEL grok-4.3 Grok model for parallel LLM calls
CLAUDE_MODEL claude-3-5-sonnet-20241022 Claude model (opt-in /analyze-image vision path; not used by default)
WIKI_OPENAI_MODEL gpt-5.5 OpenAI model for the llm-wiki response mode
MEMORY_BACKEND pg Hierarchical memory backend selector
MEMORY_SUMMARY_MODEL gpt-4o-mini Model used for session consolidation (headline + summary)
LLM_TIMEOUT 120 Per-call LLM timeout in seconds
LLM_MAX_RETRIES 2 Retry count for the reasoning LLM
AGENTIC_MAX_ITERATIONS 2 Max RAG retrieval retry limit

Consolidation cadence is a code constant, not an env var: CONSOLIDATE_EVERY = 4 in db/memory/constants.py.


Project Structure

SecondBrain_AgenticRAG/
├── docker-compose.yml           # Service orchestration
├── reset_and_run.sh             # Clean start + Debug mode
├── docs/
│   ├── architecture.svg         # System architecture diagram
│   ├── workflow.svg             # RAG workflow (incl. Classification)
│   ├── agent-flow.svg           # Detailed agent flow (incl. Classifier)
│   ├── agentic-flow.svg         # LangGraph flow-control diagram (this README)
│   └── hierarchical-memory-plan/ # Memory design + as-implemented notes
└── services/
    ├── init.sql                 # Database schema
    ├── api-service/             # FastAPI + LangGraph
    │   ├── main.py              # API endpoints
    │   ├── ask_flow.py          # Shared /ask flow + post-turn consolidation
    │   ├── agentic/
    │   │   ├── graph.py         # Workflow nodes & routing logic
    │   │   ├── state.py         # ConversationState
    │   │   ├── prompts.py       # Classification / grounding prompts
    │   │   └── tools.py         # Legacy tools (unwired)
    │   ├── db/
    │   │   └── memory/          # manager, retrieval, consolidation, lifecycle, constants
    │   ├── tools/               # Summarizer, Notion push, vocabulary, backfill_consolidate
    │   └── auth/                # Authentication middleware
    ├── frontend-service/        # Web UI (static HTML/JS)
    └── vector-service/          # Embeddings + Qdrant proxy

Roadmap & TODO

Core Improvements

  • Image Session Integration: Apply the session classifier logic to /analyze-image queries to maintain context consistency when sessions start with an image.
  • Advanced Graph Traversal: Enhance knowledge_graph_tool to explore deeper relationships in vector space.

License

MIT

About

Modernized Second Brain that keeps conversational memory, routes queries through an agentic workflow, and stores searchable knowledge across text and images.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages