A modernized Second Brain system that keeps conversational memory, routes queries through an agentic workflow, and stores searchable knowledge across text and images.
- Docker & Docker Compose
- API keys for: OpenAI, Grok (optional)
- Notion token (optional, for Notion push feature)
-
Clone the repository:
git clone https://github.com/yourusername/SecondBrian_AgenticRAG.git cd SecondBrian_AgenticRAG -
Create
.envfile with your credentials (see.env.example). -
Start all services:
docker compose up -d
-
Access the UI at http://localhost:8000
Use the provided reset script to perform a clean deployment.
Note: This script automatically enables DEBUG logging for all containers.
./reset_and_run.sh- Agentic RAG — LangGraph-powered workflow that classifies, routes, retrieves, and refines
- Hierarchical Memory (L1/L2/L3) — Working turns, episodic session headlines/summaries, and a durable semantic profile, all in PostgreSQL
- Session Consolidation — Conversations are automatically summarized into headlines every few turns and at session end, so past sessions stay searchable
- Grounded Answers — Final synthesis treats Retrieved Context as the sole source of truth about past conversations; model opinions are advisory
- Multi-LLM Parallel Calls — Queries OpenAI and Grok simultaneously with identical context
- Notion Push — Export conversation summaries directly to your Notion database (
/tools/notion_push)
| Service | Port | Override (env var) | Description |
|---|---|---|---|
| Frontend | 8000 | FRONTEND_PORT |
Web UI for chat interface — open http://localhost:8000 |
| API | 8001 | API_PORT |
FastAPI + LangGraph agentic pipeline (health: /health) |
| Vector | 8002 | VECTOR_PORT |
Embedding generation and semantic search (health: /collection/stats) |
| Qdrant | 6333 | QDRANT_HTTP_PORT |
Vector database HTTP API |
| PostgreSQL | 5432 | POSTGRES_PORT |
Sessions, messages, and hierarchical memory |
All ports are published by docker-compose.yml; set the env vars (e.g. in .env) to remap them.
Accessing from another device: the browser derives the API URL from the page's host (
http://<host>:8001), so openinghttp://<server-ip>:8000from a LAN device works out of the box.API_BASE_URLin.envonly needs changing if the API is exposed on a different host/port than<frontend-host>:8001; alocalhostvalue is ignored by non-localhost browsers.
The agentic RAG pipeline follows these stages (nodes in create_agentic_rag_graph()):
- Classify Session (
classify_session) — LLM categorizes the session as "RAG" (History Search) or "NO_RAG" (General). NO_RAG skips straight to the model fan-out. - Route (
route_query) — LLM decides whether retrieval is needed (only for RAG mode). - Retrieve (
retrieve_context) — Calls the hierarchical memory facade (MemoryManager.retrieve), gathering profile facts (L3), headline shortlist (L2), and recent working turns (L1). No LLM used. - Refine (
refine_query) — If retrieval came back empty and iterations remain, an LLM rewrites the query and retrieval is retried. - Select & Unfold (
select_and_unfold) — Expands shortlisted session headlines into full summary bodies before synthesis. - Call LLMs (
call_llms) — Parallel calls to OpenAI/Grok with the injected context. - Answer (
generate_answer) — Synthesizes the final response, grounded in the Retrieved Context.
Retrieval goes through a single facade — MemoryManager.retrieve (services/api-service/db/memory/manager.py) — which returns three tiers:
| Tier | Source | Storage | Scope | Purpose |
|---|---|---|---|---|
| L1 Working | conversation_messages |
PostgreSQL | Current session | Last 10 turns, oldest-first, for immediate context |
| L2 Episodic | conversation_sessions headlines + summaries |
PostgreSQL | Cross-session | Full-text shortlist of past-session headlines (current session excluded) |
| L3 Semantic | user_memory profile facts |
PostgreSQL | Durable, per-user | Identity/preference facts, salience-ranked under a token budget |
How the tiers are populated and searched (see docs/hierarchical-memory-plan/ for the full design):
- Consolidation lifecycle — After each assistant turn,
/askfires a debounced, fire-and-forget consolidation (ask_flow.py):consolidate_sessionno-ops until the session's message counter reachesCONSOLIDATE_EVERY(4, a constant indb/memory/constants.py), then writes a headline + summary usingMEMORY_SUMMARY_MODEL(defaultgpt-4o-mini), with a naive text fallback so L2 never stays empty.POST /sessions/newforce-consolidates the previous session before deactivating it — end-of-session is always a consolidation point. - Backfill — Consolidate every legacy session whose headline is missing:
docker compose run --rm --no-deps api python -m tools.backfill_consolidate
- Three-tier headline shortlist (L2) —
shortlist_headlines(db/memory/retrieval.py) fills up to 8 slots: (1) strict full-text AND match; (2) the same terms OR-relaxed so filler words don't cause misses; (3) rawconversation_messages.contentmatch returning the session's headline — bridging paraphrase gaps (a "cities" question can find a session whose summary never says "cities"). - Debugging — With
LOG_LEVEL=DEBUG(e.g. via./reset_and_run.sh), every tier access logs as[L1 working]/[L2 episodic]/[L3 semantic]:docker compose logs -f api | grep '\[L'. - Known gap —
promote_to_profile(L3 fact extraction) has no production trigger yet; wiring it into consolidation is a follow-up.
All LLMs receive the full retrieved context injected into their prompts:
- OpenAI —
gpt-4ofor agent reasoning (classify/route/refine/synthesize);OPENAI_MODELfor parallel calls (config defaultgpt-5-mini; note the graph's env fallback isgpt-4o-miniwhen the variable is unset) - Grok (
GROK_MODEL, defaultgrok-4.3) — Alternative perspective
/analyze-imageuses GPT (OpenAI vision) by default; the Claude/Gemini/Grok vision paths only run when explicitly requested viaselected_models. None of them are wired into the text-chat model fan-out.
This project implements an Agentic RAG (Retrieval-Augmented Generation) system powered by LangGraph, featuring intelligent routing, hierarchical memory retrieval, parallel LLM calls, and grounded synthesis.
The agentic workflow operates as a state machine with the following stages:
- Purpose: Determines the intent of a new session based on the first query.
- LLM: GPT-4o
- Modes:
RAG: History Search / Context Dependent (triggers full pipeline)NO_RAG: New Knowledge / General Assistance (bypasses retrieval for speed)
- Persistence: Remembers classification for the duration of the session.
- Purpose: Intelligently decides whether the query requires context retrieval
- LLM: GPT-4o at low temperature (for consistent reasoning)
- Logic:
- Analyzes the user's query to determine if it needs historical context
- Automatically forces retrieval for multi-turn conversations to maintain context
- Outputs:
should_retrievedecision flag
- Purpose: Gathers relevant information via the hierarchical memory facade
- No LLM used — Pure data retrieval for speed and cost efficiency
- Implementation: acquires a DB connection and calls
MemoryManager.retrieve(conn, query, user_id, session_id), which returns:- Profile facts (L3,
user_memory) — durable cross-session identity/preferences - Headlines (L2,
conversation_sessions) — shortlisted candidate summaries from other sessions - Working turns (L1,
conversation_messages) — recent turns from the current session
- Profile facts (L3,
- Each item is tagged with its
source(profile/headlines/working) inretrieved_context.
- Purpose: Improves the query if initial retrieval returns insufficient context
- Trigger: Activated when
retrieved_contextis empty but iterations remain (AGENTIC_MAX_ITERATIONS) - LLM: GPT-4o refines the query phrasing
- Iteration: Retries retrieval with the refined query
- Purpose: Expands shortlisted L2 headlines into full summary bodies before synthesis
- Behavior: Collects
headlinesitems fromretrieved_context, unfolds their bodies, and appends them asunfoldedcontext items - Crash-safe: If no headlines were retrieved (or unfolding fails), it is a pass-through that leaves
retrieved_contextuntouched
- Purpose: Parallel fan-out to multiple LLM providers with injected context
- Models: OpenAI (
OPENAI_MODEL), Grok (GROK_MODEL) - Context Injection:
- Retrieved context formatted into a structured prompt
- Session history included for conversational coherence
- Each model receives identical context for consistent comparison
- Execution: All models called simultaneously via
asyncio.gather(); per-model answers and errors tracked separately (llm_answers/llm_errors)
- Purpose: Synthesizes a final, grounded answer
- LLM: GPT-4o as the synthesis engine
- Grounding rules (
agentic/prompts.py):- The Retrieved Context is the only source of truth about past conversations — if model opinions conflict with it, or claim past discussions that do not appear in it, the context wins
- If a topic is absent from the Retrieved Context, the answer states plainly that there is no record of it
- Input: Retrieved (and unfolded) context, individual answers from all LLMs, and conversation history
The workflow uses LangGraph's StateGraph with typed state (ConversationState, agentic/state.py):
{
"session_id": str, # Current conversation session
"user_id": str, # User identifier
"session_mode": str, # "RAG" or "NO_RAG" classification
"current_query": str, # User's question (can be refined)
"messages": List[dict], # Conversation history (reducer-annotated: operator.add)
"should_retrieve": bool, # Routing decision
"should_refine": bool, # Whether to refine the query
"retrieved_context": List, # Context items tagged by memory tier/source
"selected_models": List, # Which LLMs to call
"llm_answers": Dict, # Individual model responses
"llm_errors": Dict, # Error tracking per model (not fed to synthesis)
"final_answer": str, # Synthesized response
"response_mode": str, # "fast" or "llm-wiki"
"skill_prompt": str, # Active skill instructions (llm-wiki mode)
"agent_thoughts": List, # Debugging/transparency log
"iteration_count": int, # Current iteration
"max_iterations": int, # Iteration limit (AGENTIC_MAX_ITERATIONS, default 2)
"flow_timing": Dict # Per-node / per-model timing metadata
}| Feature | Description |
|---|---|
| Session Classification | Distinguishes between historical search vs general requests on first query |
| Conditional Routing | Not all queries need retrieval—simple questions go straight to the model fan-out |
| Hierarchical Memory (L1/L2/L3) | Working turns, episodic session headlines/summaries, and semantic profile facts through one facade |
| Session Consolidation | Debounced headline+summary generation per turn, forced at session end, backfillable for legacy sessions |
| Grounded Synthesis | Retrieved Context is the sole source of truth about past conversations; model opinions are advisory |
| Parallel LLM Calls | Calls OpenAI and Grok simultaneously for diverse perspectives |
| Transparent Logging | agent_thoughts and flow_timing track decisions and per-node latency for debugging |
The workflow uses conditional edges for dynamic routing:
Edges: classify_session -> call_llms (NO_RAG) | route_query; route_query -> retrieve_context | call_llms | END; retrieve_context -> refine_query | select_and_unfold; refine_query -> retrieve_context; select_and_unfold -> call_llms; call_llms -> generate_answer -> END.
Traditional RAG systems follow a fixed pipeline: retrieve → generate. This agentic approach adds:
- Intelligence: Decides when retrieval is necessary
- Adaptability: Refines queries if initial retrieval fails
- Redundancy: Calls multiple LLMs for robustness
- Context-Awareness: Full conversation memory across sessions
- Observability: Explicit state tracking and decision logging
This architecture ensures that answers are not just generated from context, but intelligently orchestrated through a reasoning process that adapts to each query's needs.
All routes live in services/api-service/main.py unless noted.
| Method | Endpoint | Description |
|---|---|---|
GET |
/health |
Service health check |
POST |
/sessions/create |
Create a new conversation session |
GET |
/sessions |
List user sessions |
GET |
/sessions/{session_id} |
Get session detail (incl. messages) |
DELETE |
/sessions/{session_id} |
Delete a session |
PUT |
/sessions/{session_id}/title |
Rename a session |
POST |
/sessions/new |
Start a new session (consolidates + deactivates the previous one) |
GET |
/sessions/{session_id}/wiki-export.zip |
Export session as wiki bundle (zip) |
GET |
/sessions/{session_id}/wiki-export.md |
Export session as a single markdown file |
| Method | Endpoint | Description |
|---|---|---|
POST |
/ask |
Main question endpoint (triggers agentic RAG) |
POST |
/analyze-image |
Vision analysis (GPT by default; Claude/Gemini/Grok via explicit selected_models) |
POST |
/reset |
Reset a session's working memory |
| Method | Endpoint | Description |
|---|---|---|
POST |
/query |
Simple query endpoint (stub implementation) |
POST |
/store-knowledge |
Store a Q&A pair in the vector database |
GET |
/search/{query} |
Search stored knowledge (configurable threshold) |
GET |
/status |
Status of all microservices |
| Method | Endpoint | Description |
|---|---|---|
POST |
/vector/search |
Semantic search |
POST |
/vector/topics |
Topics extracted from actual Q&A sessions |
POST |
/vector/stats |
Vector collection stats |
GET |
/vector/knowledge-graph |
User's knowledge graph |
POST |
/vector/by-topic |
User's entries filtered by topic |
| Method | Endpoint | Description |
|---|---|---|
POST |
/tools/daily_summarizer |
Generate daily summary (tools/summarizer.py) |
POST |
/tools/notion_push |
Push a summary to Notion (tools/notion_push.py) |
POST |
/tools/daily_summarize_and_push |
Summarize + Notion push combo (tools/daily_summarize_and_push.py) |
| — | /vocabulary/* |
Vocabulary memory CRUD, search, spaced-repetition review, and stats (13 routes, tools/vocabulary_routes.py) |
agentic/tools.py still defines semantic_search_tool, session_history_tool, user_memories_tool, and knowledge_graph_tool, but the graph no longer uses them — retrieval goes through MemoryManager. They are kept for reference and slated for consolidation.
| Variable | Default | Description |
|---|---|---|
OPENAI_API_KEY |
— | OpenAI API key (required) |
LOG_LEVEL |
INFO |
Global log level (DEBUG enables [L1]/[L2]/[L3] memory tracing) |
OPENAI_MODEL |
gpt-5-mini |
OpenAI model for parallel LLM calls (config default; graph falls back to gpt-4o-mini if unset) |
GEMINI_MODEL |
gemini-2.0-flash-exp |
Gemini model (summarization path + opt-in vision analysis; not in the ask fan-out) |
GROK_MODEL |
grok-4.3 |
Grok model for parallel LLM calls |
CLAUDE_MODEL |
claude-3-5-sonnet-20241022 |
Claude model (opt-in /analyze-image vision path; not used by default) |
WIKI_OPENAI_MODEL |
gpt-5.5 |
OpenAI model for the llm-wiki response mode |
MEMORY_BACKEND |
pg |
Hierarchical memory backend selector |
MEMORY_SUMMARY_MODEL |
gpt-4o-mini |
Model used for session consolidation (headline + summary) |
LLM_TIMEOUT |
120 |
Per-call LLM timeout in seconds |
LLM_MAX_RETRIES |
2 |
Retry count for the reasoning LLM |
AGENTIC_MAX_ITERATIONS |
2 |
Max RAG retrieval retry limit |
Consolidation cadence is a code constant, not an env var:
CONSOLIDATE_EVERY = 4indb/memory/constants.py.
SecondBrain_AgenticRAG/
├── docker-compose.yml # Service orchestration
├── reset_and_run.sh # Clean start + Debug mode
├── docs/
│ ├── architecture.svg # System architecture diagram
│ ├── workflow.svg # RAG workflow (incl. Classification)
│ ├── agent-flow.svg # Detailed agent flow (incl. Classifier)
│ ├── agentic-flow.svg # LangGraph flow-control diagram (this README)
│ └── hierarchical-memory-plan/ # Memory design + as-implemented notes
└── services/
├── init.sql # Database schema
├── api-service/ # FastAPI + LangGraph
│ ├── main.py # API endpoints
│ ├── ask_flow.py # Shared /ask flow + post-turn consolidation
│ ├── agentic/
│ │ ├── graph.py # Workflow nodes & routing logic
│ │ ├── state.py # ConversationState
│ │ ├── prompts.py # Classification / grounding prompts
│ │ └── tools.py # Legacy tools (unwired)
│ ├── db/
│ │ └── memory/ # manager, retrieval, consolidation, lifecycle, constants
│ ├── tools/ # Summarizer, Notion push, vocabulary, backfill_consolidate
│ └── auth/ # Authentication middleware
├── frontend-service/ # Web UI (static HTML/JS)
└── vector-service/ # Embeddings + Qdrant proxy
- Image Session Integration: Apply the session classifier logic to
/analyze-imagequeries to maintain context consistency when sessions start with an image. - Advanced Graph Traversal: Enhance
knowledge_graph_toolto explore deeper relationships in vector space.
MIT