Skip to content

About

A practical open-source guide to designing production AI systems: RAG, agents, evals, observability, safety, cost, latency, and architecture tradeoffs.

Resources

Contributing

Stars

11 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AI System Design

AI System Design hero

A practical, open-source course and handbook for designing production AI systems.

Most system design resources teach databases, caches, queues, APIs, replication, sharding, and scaling. Those still matter. But AI-native products add new failure modes and design decisions:

  • Probabilistic model behavior
  • Hallucinations and unverifiable answers
  • Retrieval quality and stale knowledge
  • Prompt and model version drift
  • Tool misuse and agentic failure loops
  • Evaluation pipelines and regression testing
  • Latency budgets shaped by model calls and token volume
  • Cost controls across prompts, retrieval, reranking, and inference
  • Security risks such as prompt injection, data leakage, and unsafe tool execution

This repo exists to map those problems clearly and turn them into a rigorous learning path.

What This Is

AI System Design is a course-quality field guide for engineers designing real AI products. It focuses on architecture, tradeoffs, failure modes, evaluation, observability, security, and production constraints.

The goal is not to collect AI news. The goal is to explain how to make system design decisions when the core component is a model whose output is useful but not guaranteed.

Who This Is For

  • Backend engineers moving into AI products
  • ML engineers who need product and infrastructure architecture context
  • Founders building AI-native SaaS products
  • Senior engineers preparing for AI system design interviews
  • Devtool and platform engineers building internal AI infrastructure

What This Is Not

  • Not an AI news feed
  • Not a prompt-hack collection
  • Not vendor marketing
  • Not a generic ML theory course
  • Not an "awesome links" repository
  • Not a place for unsourced claims or shallow summaries

Start Here

  1. Course
  2. Syllabus
  3. Glossary
  4. What Is AI System Design?
  5. RAG System Design
  6. Agent Tool-Use System Design
  7. Evaluation Pipeline Pattern
  8. RAG vs Fine-Tuning

Build From Scratch

Every core component of a production AI system, implemented in pure Python standard library. No frameworks, no API keys, no network. Each module runs with one command, validates itself with asserts, and ends with what production systems do differently.

The Atlas

The complete map of AI system design as a field: atlas/README.md. Ten territories, every topic in one tight paragraph, each marked as covered with a link into this repo or honestly marked as planned. Use it to find what to read next, or what to contribute.

Hands-On Labs

  1. RAG Retrieval Eval Lab
  2. Eval Set Runner Lab
  3. Tool Policy Simulator Lab
  4. Structured Output Validator Lab

High-Value Pages

Course Completion

The full path includes readings, labs, assignments, review guides, and a capstone:

Diagrams And Sources

The architecture diagrams in this repo are original Mermaid diagrams written for the course. External references are used as sources for claims, standards, and production guidance; they are not copied as diagrams or visual assets.

Start with the source map for the current evidence base.

Repo Map

ai-system-design/
├── foundations/             # Core concepts and mental models
├── patterns/                # Reusable architecture patterns
├── decision-guides/         # Tradeoff-driven engineering decisions
├── case-studies/            # Realistic system design walkthroughs
├── labs/                    # Hands-on local exercises
├── assignments/             # Design problems and rubrics
├── answer-keys/             # Review guides for assignments
├── capstones/               # Final end-to-end design projects
├── design-reviews/          # Worked architecture reviews
├── templates/               # Design doc, capstone, and frontier note templates
├── evals-observability/     # Testing, tracing, monitoring, and feedback loops
├── security/                # AI-specific threat models and mitigations
├── reference-architectures/ # Production-ready blueprints
├── frontier-notes/          # Cutting-edge changes translated into production impact
└── resources/               # Curated source map, not a dumping ground

Content Standard

Every serious page should answer:

  • What problem does this solve?
  • When should you use it?
  • What is the architecture?
  • What are the core components?
  • What are the tradeoffs?
  • What fails in production?
  • How do you evaluate it?
  • What should you observe?
  • What are the cost and latency implications?
  • What are the security risks?
  • What sources support the claims?

See CONTENT_STANDARD.md.

Contributing

This repo should be useful because it is selective. Contributions are welcome, but the bar is intentionally high.

Good contributions include:

  • Production architecture patterns
  • Clear decision guides
  • Case studies with concrete tradeoffs
  • Failure modes from real systems
  • Evaluation and observability methods
  • Primary-source-backed frontier notes

Start with CONTRIBUTING.md.

License

Content is licensed under CC BY 4.0 unless otherwise noted.

About

A practical open-source guide to designing production AI systems: RAG, agents, evals, observability, safety, cost, latency, and architecture tradeoffs.

Resources

Contributing

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages