Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 71 additions & 0 deletions .github/workflows/smoke.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
name: Streamlit smoke tests

on:
push:
branches: [main]
pull_request:
workflow_dispatch:

permissions:
contents: read

jobs:
smoke:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v4

- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
cache-dependency-path: |
requirements.txt
requirements-dev.txt

- name: Install Python dependencies
run: pip install -r requirements-dev.txt

- name: Install Playwright Chromium
run: python -m playwright install --with-deps chromium

- name: Run Streamlit + Playwright smoke tests
run: python -m pytest tests/smoke -v

openai-provider-smoke:
name: OpenAI provider smoke (optional, OPENAI_API_KEY)
runs-on: ubuntu-latest
timeout-minutes: 10
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
steps:
- name: Check for OPENAI_API_KEY
id: key
run: |
if [ -n "$OPENAI_API_KEY" ]; then
echo "present=true" >> "$GITHUB_OUTPUT"
else
echo "present=false" >> "$GITHUB_OUTPUT"
echo "::notice title=OpenAI provider smoke skipped::OPENAI_API_KEY is not available to this run (not configured, or a fork or Dependabot PR). Run this workflow manually to check the live OpenAI path."
fi

- if: steps.key.outputs.present == 'true'
uses: actions/checkout@v4

- if: steps.key.outputs.present == 'true'
uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
cache-dependency-path: |
requirements.txt
requirements-dev.txt

- name: Install Python dependencies
if: steps.key.outputs.present == 'true'
run: pip install -r requirements-dev.txt

- name: Run live OpenAI provider smoke tests
if: steps.key.outputs.present == 'true'
run: python -m pytest tests/provider -v -rs
84 changes: 84 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
# AGENTS.md

Notes for contributors and coding agents working on this repo.

## Layout

- `chat_with_pdf_query.py`: default app. Uses `CouchbaseQueryVectorStore` (Query + Index services, Hyperscale/Composite vector indexes).
- `chat_with_pdf.py`: uses `CouchbaseSearchVectorStore` (Search service, needs `INDEX_NAME`).
- `tests/smoke/`: Streamlit + Playwright smoke tests for both entrypoints (no credentials).
- `tests/provider/`: optional live OpenAI provider smoke test (needs `OPENAI_API_KEY`, no Couchbase).
- `.github/workflows/smoke.yml`: runs the smoke tests on every PR and push to `main`, plus the optional provider job.

## Test tiers

| Tier | What | When | Required? |
| --- | --- | --- | --- |
| 1 | `tests/smoke` Streamlit + Playwright smoke, no secrets | Every PR, push to `main`, manual dispatch | Yes. The required PR check |
| 2 | `tests/provider` live OpenAI provider smoke | CI job `openai-provider-smoke` when `OPENAI_API_KEY` is available; skips with a notice otherwise | No. Optional signal |
| 3 | Full Couchbase-backed RAG run (checklist below) | Manually, by a maintainer, before merging dependency updates | Yes. Evidence goes in the PR |

Tiers 1 and 2 do not replace tier 3.

## Install

```bash
pip install -r requirements-dev.txt # app deps + pytest + playwright
python -m playwright install --with-deps chromium
```

`requirements.txt` holds only the app's runtime dependencies. Dependabot bumps both files.

## Smoke tests (CI, no credentials)

```bash
python -m pytest tests/smoke -v
```

These tests need no Couchbase cluster and no OpenAI key. They start each app headless on a free port, open it in headless Chromium, and fail if a Python traceback or `st.exception` appears in the page or the server log. For each entrypoint they check:

1. **Unconfigured start**: the real app with no settings loads, imports every real dependency, and shows the `OPENAI_API_KEY environment variable is not set` message.
2. **Auth gate**: with `AUTH_ENABLED=True`, the password form renders and rejects a wrong password.
3. **Stubbed full flow**: `tests/smoke/stubbed_app.py` runs the app with the Couchbase cluster, vector store, and cache, plus the OpenAI embeddings and chat model, swapped for in-memory LangChain fakes. The test checks that the expected controls render, uploads a small generated PDF (real `pypdf` loader and text splitter), asks a question, and asserts that the RAG answer used the retrieved PDF context and the pure-LLM answer did not.

The tests set placeholder values for every app setting and use an isolated `HOME`, so a local `.streamlit/secrets.toml` or real credentials in your shell are never used.

**What the smoke tests do not cover:** real Couchbase connectivity, vector index definitions, `langchain-couchbase` query/search behaviour against a cluster, OpenAI API compatibility (model names, request shapes), and the LLM cache against a real collection. The optional provider smoke covers the OpenAI part. The rest is why live validation stays manual (below).

## OpenAI provider smoke (optional, needs `OPENAI_API_KEY`)

```bash
OPENAI_API_KEY=... python -m pytest tests/provider -v -rs
```

This calls the real OpenAI API and does not touch Couchbase. It reads the `OpenAIEmbeddings(...)` and `ChatOpenAI(...)` arguments from both app entrypoints, so it checks the exact models and options the apps use:

- **Embeddings**: `embed_query` returns a 1536-dim vector of finite floats (the README index definitions use 1536 dims).
- **Chat**: each distinct `ChatOpenAI` configuration returns a non-empty reply to a one-line prompt.

Without `OPENAI_API_KEY` the tests are skipped with a clear reason. With an invalid key or a retired model they fail. In CI the `openai-provider-smoke` job uses the `OPENAI_API_KEY` Actions secret. Secrets are not passed to fork or Dependabot PRs, so there the job posts a "skipped" notice and passes. To check the live OpenAI path for such a PR, run the workflow manually (**Actions → Streamlit smoke tests → Run workflow**) on the PR branch.

## Manual live validation for dependency-update PRs

A green smoke run is necessary but not sufficient to merge a dependency bump (for example `langchain-*`, `couchbase`, `streamlit`, `pypdf`). Before merging, a maintainer runs the app against a real cluster and records this evidence in the PR:

1. **App boot**: `streamlit run chat_with_pdf_query.py` (and `streamlit run chat_with_pdf.py` if the bump touches `langchain-couchbase`, `couchbase`, or Search) starts with no errors, and the installed versions match the PR (`pip freeze | grep -iE "langchain|couchbase|streamlit|pypdf|openai"`).
2. **Couchbase / vector index state**: the cluster version, the bucket/scope/collection used, and that the index exists. For the Query app that is the Hyperscale/Composite vector index from the README's "Index Configuration" section. For the Search app it is the Search index named by `INDEX_NAME` (see "Create the Search Index").
3. **Provider credentials**: confirm that `OPENAI_API_KEY` was available and that the configured models were accessible. State only that it was set, never the value.
4. **Sample input / query**: the PDF you uploaded (name and approximate size) and the question you asked.
5. **Observed response**: the RAG answer and the pure-LLM answer. Ask the same question a second time to confirm the cached response is returned.
6. **Screenshots or traces** of the upload confirmation and the chat answers, when practical.

## Required settings / secret names

Set these in `.streamlit/secrets.toml` (copy `.streamlit/secrets.example.toml`) or as environment variables. Never commit real values, and never paste them into PRs, issues, or logs.

| Name | Used by |
| --- | --- |
| `OPENAI_API_KEY` | both apps |
| `DB_CONN_STR`, `DB_USERNAME`, `DB_PASSWORD` | both apps |
| `DB_BUCKET`, `DB_SCOPE`, `DB_COLLECTION`, `CACHE_COLLECTION` | both apps |
| `INDEX_NAME` | `chat_with_pdf.py` only |
| `AUTH_ENABLED`, `LOGIN_PASSWORD` | optional app password gate |

The required CI smoke job needs **no** repository secrets. The optional `openai-provider-smoke` job uses the `OPENAI_API_KEY` Actions secret when it is available.
18 changes: 18 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,24 @@ All LLM responses are cached in the collection specified. If the same exact ques

`pip install -r requirements.txt`

### Smoke tests

Streamlit + Playwright smoke tests for both apps run in CI without any credentials:

```bash
pip install -r requirements-dev.txt
python -m playwright install --with-deps chromium
python -m pytest tests/smoke -v
```

An optional live OpenAI provider check (no Couchbase) runs when `OPENAI_API_KEY` is set and is skipped otherwise:

```bash
python -m pytest tests/provider -v -rs
```

See [AGENTS.md](AGENTS.md) for what each tier covers and the manual live-validation checklist for dependency updates. Neither test suite replaces a full Couchbase-backed RAG run.

### Set the environment secrets

Copy the `secrets.example.toml` file in `.streamlit` folder and rename it to `secrets.toml` and replace the placeholders with the actual values for your environment.
Expand Down
5 changes: 3 additions & 2 deletions chat_with_pdf.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,8 +40,9 @@ def save_to_vector_store(uploaded_file, vector_store):

with open(temp_file_path, "wb") as f:
f.write(uploaded_file.getvalue())
loader = PyPDFLoader(temp_file_path)
docs = loader.load()

loader = PyPDFLoader(temp_file_path)
docs = loader.load()

text_splitter = RecursiveCharacterTextSplitter(
chunk_size=1500, chunk_overlap=150
Expand Down
5 changes: 3 additions & 2 deletions chat_with_pdf_query.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,8 +41,9 @@ def save_to_vector_store(uploaded_file, vector_store):

with open(temp_file_path, "wb") as f:
f.write(uploaded_file.getvalue())
loader = PyPDFLoader(temp_file_path)
docs = loader.load()

loader = PyPDFLoader(temp_file_path)
docs = loader.load()

text_splitter = RecursiveCharacterTextSplitter(
chunk_size=1500, chunk_overlap=150
Expand Down
3 changes: 3 additions & 0 deletions requirements-dev.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
-r requirements.txt
pytest==9.1.1
playwright==1.63.0
62 changes: 62 additions & 0 deletions tests/provider/test_openai_provider.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
"""Optional live OpenAI provider smoke test. Needs OPENAI_API_KEY; no Couchbase.

The embedding and chat model settings are read from both app entrypoints, so
this checks the exact models and ChatOpenAI arguments the apps use.
"""

import ast
import math
import os
from pathlib import Path

import pytest

REPO_ROOT = Path(__file__).resolve().parents[2]
APPS = ["chat_with_pdf_query.py", "chat_with_pdf.py"]

# Both README index definitions use a 1536-dim vector field.
EMBEDDING_DIMS = 1536

pytestmark = pytest.mark.skipif(
not os.environ.get("OPENAI_API_KEY"),
reason="OPENAI_API_KEY is not set; skipping live OpenAI provider smoke test",
)


def _constructor_calls(name):
"""Literal keyword arguments of every `name(...)` call in the apps."""
calls = set()
for app in APPS:
for node in ast.walk(ast.parse((REPO_ROOT / app).read_text())):
if isinstance(node, ast.Call) and isinstance(node.func, ast.Name) and node.func.id == name:
kwargs = {kw.arg: ast.literal_eval(kw.value) for kw in node.keywords}
calls.add(tuple(sorted(kwargs.items())))
assert calls, f"no {name}(...) call found in {APPS}"
return [dict(c) for c in sorted(calls)]


def _ids(calls):
return [",".join(f"{k}={v}" for k, v in c.items()) or "defaults" for c in calls]


EMBEDDING_CALLS = _constructor_calls("OpenAIEmbeddings")
CHAT_CALLS = _constructor_calls("ChatOpenAI")


@pytest.mark.parametrize("kwargs", EMBEDDING_CALLS, ids=_ids(EMBEDDING_CALLS))
def test_embeddings_return_numeric_vector(kwargs):
from langchain_openai import OpenAIEmbeddings

vector = OpenAIEmbeddings(**kwargs).embed_query("Couchbase vector search")
assert len(vector) == EMBEDDING_DIMS
assert all(isinstance(x, float) and math.isfinite(x) for x in vector)
assert any(x != 0 for x in vector)


@pytest.mark.parametrize("kwargs", CHAT_CALLS, ids=_ids(CHAT_CALLS))
def test_chat_model_returns_text(kwargs):
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(**kwargs, max_retries=1)
reply = llm.invoke("Reply with the single word: pong")
assert isinstance(reply.content, str) and reply.content.strip()
Loading
Loading