Skip to content

Add Streamlit + Playwright smoke tests and CI workflow - #156

Open
dex-the-ai wants to merge 3 commits into
mainfrom
dex/streamlit-playwright-smoke
Open

dex-the-ai wants to merge 3 commits into
mainfrom
dex/streamlit-playwright-smoke

Conversation

@dex-the-ai

@dex-the-ai dex-the-ai commented Oct 8, 2026 •

Copy link
Copy Markdown

Summary

Adds a lightweight Streamlit + Playwright smoke-test harness, so dependency-update PRs (Dependabot langchain-*, couchbase, streamlit, pypdf) get a real CI check before manual RAG validation.

Kanban task: t_75db56c0 (board couchbase-examples): "Add Streamlit Playwright smoke tests for couchbase-examples/rag-demo"

What's covered (no Couchbase or OpenAI credentials needed)

Each test runs against both chat_with_pdf_query.py and chat_with_pdf.py:

Test What it proves
test_unconfigured_app_reports_missing_settings The real app imports every real dependency, boots headless, and shows the missing-OPENAI_API_KEY message instead of crashing
test_auth_gate_renders With AUTH_ENABLED=True, the password form renders and rejects a wrong password
test_stubbed_app_upload_and_chat Full UI with the Couchbase cluster, vector store, and cache, plus OpenAI embeddings and chat, swapped for in-memory LangChain fakes (tests/smoke/stubbed_app.py). Checks that the expected controls render, uploads a generated PDF (real pypdf + text splitter), asks a question, and asserts that the RAG answer used the retrieved PDF context and the pure-LLM answer did not

Every test fails on a Streamlit stException element, on visible Traceback text, or on a traceback in the Streamlit server log.

Bug fix found by the smoke test

save_to_vector_store called PyPDFLoader inside the with open(...) block, before the buffered write was flushed. Any PDF smaller than about 8 KB was read as empty and raised pypdf.errors.EmptyFileError: Cannot read an empty file. I reproduced this outside the harness: a 580-byte PDF fails inside the with and loads fine after it closes. The fix (commit de14220) dedents the two loader lines in both apps.

Optional OpenAI provider smoke (added in 00b6a58)

tests/provider/test_openai_provider.py calls the real OpenAI API and does not use Couchbase. It reads the OpenAIEmbeddings(...) and ChatOpenAI(...) arguments from both entrypoints, so it checks the exact models and options the apps use:

  • OpenAIEmbeddings().embed_query(...) returns a 1536-dim vector of finite floats (this matches the README index dims).
  • Each distinct ChatOpenAI config (gpt-5.4-nano, streaming, with and without temperature=0) returns a non-empty reply to a minimal prompt.

The tests skip with a clear reason when OPENAI_API_KEY is unset, and fail on an invalid key or a retired model. The new openai-provider-smoke CI job passes the OPENAI_API_KEY Actions secret. When the secret is unavailable (forks, Dependabot PRs), it posts a "skipped" notice and passes; maintainers can use Run workflow to check those PRs. The no-secret smoke job is still the required check.

Test tiers

Tier What Required?
1 tests/smoke: Streamlit + Playwright, no secrets Yes. Required PR check
2 tests/provider: live OpenAI, no Couchbase Optional. Runs when OPENAI_API_KEY is available
3 Full Couchbase-backed RAG run (AGENTS.md checklist) Yes. Manual evidence in dependency-update PRs

Tiers 1 and 2 do not replace tier 3.

Other changes

  • .github/workflows/smoke.yml: Python 3.12, pip install -r requirements-dev.txt, playwright install --with-deps chromium, pytest tests/smoke -v. Runs on PRs, pushes to main, and manual dispatch. The smoke job needs no secrets; the optional openai-provider-smoke job uses OPENAI_API_KEY when it is available.
  • requirements-dev.txt: -r requirements.txt plus pinned pytest and playwright, so Dependabot picks these up too.
  • AGENTS.md (new): install and smoke commands, what the smoke tests do and don't cover, the manual live-validation checklist for dependency-update PRs, and required secret names only.
  • README.md: short "Smoke tests" section that points to AGENTS.md.

Live RAG validation stays manual

The smoke tests don't reach a real cluster or the OpenAI API. Per AGENTS.md, dependency-update PRs still need evidence of: app boot plus installed versions, Couchbase and vector index state, that OPENAI_API_KEY was available (never its value), the sample PDF and question, the observed RAG, pure-LLM, and cached responses, and screenshots or traces when practical.

Local test result

$ python -m pytest tests/smoke -v      # Python 3.12.13, fresh venv from requirements-dev.txt, Chromium headless shell
tests/smoke/test_streamlit_apps.py::test_unconfigured_app_reports_missing_settings[chat_with_pdf_query.py] PASSED [ 16%]
tests/smoke/test_streamlit_apps.py::test_unconfigured_app_reports_missing_settings[chat_with_pdf.py] PASSED [ 33%]
tests/smoke/test_streamlit_apps.py::test_auth_gate_renders[chat_with_pdf_query.py] PASSED [ 50%]
tests/smoke/test_streamlit_apps.py::test_auth_gate_renders[chat_with_pdf.py] PASSED [ 66%]
tests/smoke/test_streamlit_apps.py::test_stubbed_app_upload_and_chat[chat_with_pdf_query.py] PASSED [ 83%]
tests/smoke/test_streamlit_apps.py::test_stubbed_app_upload_and_chat[chat_with_pdf.py] PASSED [100%]
============================== 6 passed in 58.79s ==============================
$ python -m pytest tests/provider -v -rs     # with OPENAI_API_KEY set (value not shown)
tests/provider/test_openai_provider.py::test_embeddings_return_numeric_vector[defaults] PASSED [ 33%]
tests/provider/test_openai_provider.py::test_chat_model_returns_text[model=gpt-5.4-nano,streaming=True] PASSED [ 66%]
tests/provider/test_openai_provider.py::test_chat_model_returns_text[model=gpt-5.4-nano,streaming=True,temperature=0] PASSED [100%]
============================== 3 passed in 4.01s ===============================

$ env -u OPENAI_API_KEY python -m pytest tests/provider -v -rs
SKIPPED [1] ... OPENAI_API_KEY is not set; skipping live OpenAI provider smoke test
SKIPPED [2] ... OPENAI_API_KEY is not set; skipping live OpenAI provider smoke test
============================== 3 skipped in 0.04s ==============================

With an invalid key, all 3 provider tests fail, so a bad key can't pass by skipping. CI run 37814472981: smoke passed 6/6, and OpenAI provider smoke passed 3/3 against the live API.

Before the fix, test_stubbed_app_upload_and_chat[chat_with_pdf_query.py] failed with the EmptyFileError traceback rendered in the app. This shows the traceback detection works.

cc @nithishr @AayushTyagi1

🤖 Generated with Claude Code

dex-the-ai and others added 2 commits October 8, 2026 16:23
save_to_vector_store ran PyPDFLoader inside the `with open(...)` block,
before the buffered write was flushed. PDFs smaller than the write
buffer (~8 KB) were read as empty and raised pypdf EmptyFileError.
Found by the new Streamlit smoke test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- tests/smoke: boots both entrypoints headless and drives them with
  Playwright: unconfigured start, auth gate, and a full upload + chat
  flow with Couchbase/OpenAI swapped for in-memory LangChain fakes.
  Fails on any visible traceback or traceback in the server log.
- .github/workflows/smoke.yml: runs the smoke tests on PRs and main.
- requirements-dev.txt: pytest + playwright on top of requirements.txt.
- AGENTS.md: install/test commands, manual live-validation checklist
  for dependency-update PRs, required secret names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tests/provider calls the real OpenAI API (no Couchbase), using the
OpenAIEmbeddings/ChatOpenAI arguments parsed from both app entrypoints:
embeddings must return a 1536-dim finite float vector, and each chat
config must return non-empty text. Skips when OPENAI_API_KEY is unset.

The openai-provider-smoke job runs it when the secret is available and
posts a skip notice otherwise (forks, Dependabot). The no-secret smoke
job stays the required check. AGENTS.md documents the three tiers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant