Repository navigation
Add Streamlit + Playwright smoke tests and CI workflow - #156
Open
dex-the-ai wants to merge 3 commits into
Open
dex-the-ai wants to merge 3 commits into
dex-the-ai wants to merge 3 commits into
Conversation
save_to_vector_store ran PyPDFLoader inside the `with open(...)` block, before the buffered write was flushed. PDFs smaller than the write buffer (~8 KB) were read as empty and raised pypdf EmptyFileError. Found by the new Streamlit smoke test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- tests/smoke: boots both entrypoints headless and drives them with Playwright: unconfigured start, auth gate, and a full upload + chat flow with Couchbase/OpenAI swapped for in-memory LangChain fakes. Fails on any visible traceback or traceback in the server log. - .github/workflows/smoke.yml: runs the smoke tests on PRs and main. - requirements-dev.txt: pytest + playwright on top of requirements.txt. - AGENTS.md: install/test commands, manual live-validation checklist for dependency-update PRs, required secret names. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tests/provider calls the real OpenAI API (no Couchbase), using the OpenAIEmbeddings/ChatOpenAI arguments parsed from both app entrypoints: embeddings must return a 1536-dim finite float vector, and each chat config must return non-empty text. Skips when OPENAI_API_KEY is unset. The openai-provider-smoke job runs it when the secret is available and posts a skip notice otherwise (forks, Dependabot). The no-secret smoke job stays the required check. AGENTS.md documents the three tiers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a lightweight Streamlit + Playwright smoke-test harness, so dependency-update PRs (Dependabot
langchain-*,couchbase,streamlit,pypdf) get a real CI check before manual RAG validation.Kanban task:
t_75db56c0(boardcouchbase-examples): "Add Streamlit Playwright smoke tests for couchbase-examples/rag-demo"What's covered (no Couchbase or OpenAI credentials needed)
Each test runs against both
chat_with_pdf_query.pyandchat_with_pdf.py:test_unconfigured_app_reports_missing_settingsOPENAI_API_KEYmessage instead of crashingtest_auth_gate_rendersAUTH_ENABLED=True, the password form renders and rejects a wrong passwordtest_stubbed_app_upload_and_chattests/smoke/stubbed_app.py). Checks that the expected controls render, uploads a generated PDF (realpypdf+ text splitter), asks a question, and asserts that the RAG answer used the retrieved PDF context and the pure-LLM answer did notEvery test fails on a Streamlit
stExceptionelement, on visibleTracebacktext, or on a traceback in the Streamlit server log.Bug fix found by the smoke test
save_to_vector_storecalledPyPDFLoaderinside thewith open(...)block, before the buffered write was flushed. Any PDF smaller than about 8 KB was read as empty and raisedpypdf.errors.EmptyFileError: Cannot read an empty file. I reproduced this outside the harness: a 580-byte PDF fails inside thewithand loads fine after it closes. The fix (commit de14220) dedents the two loader lines in both apps.Optional OpenAI provider smoke (added in 00b6a58)
tests/provider/test_openai_provider.pycalls the real OpenAI API and does not use Couchbase. It reads theOpenAIEmbeddings(...)andChatOpenAI(...)arguments from both entrypoints, so it checks the exact models and options the apps use:OpenAIEmbeddings().embed_query(...)returns a 1536-dim vector of finite floats (this matches the README index dims).ChatOpenAIconfig (gpt-5.4-nano, streaming, with and withouttemperature=0) returns a non-empty reply to a minimal prompt.The tests skip with a clear reason when
OPENAI_API_KEYis unset, and fail on an invalid key or a retired model. The newopenai-provider-smokeCI job passes theOPENAI_API_KEYActions secret. When the secret is unavailable (forks, Dependabot PRs), it posts a "skipped" notice and passes; maintainers can use Run workflow to check those PRs. The no-secretsmokejob is still the required check.Test tiers
tests/smoke: Streamlit + Playwright, no secretstests/provider: live OpenAI, no CouchbaseOPENAI_API_KEYis availableTiers 1 and 2 do not replace tier 3.
Other changes
.github/workflows/smoke.yml: Python 3.12,pip install -r requirements-dev.txt,playwright install --with-deps chromium,pytest tests/smoke -v. Runs on PRs, pushes tomain, and manual dispatch. Thesmokejob needs no secrets; the optionalopenai-provider-smokejob usesOPENAI_API_KEYwhen it is available.requirements-dev.txt:-r requirements.txtplus pinnedpytestandplaywright, so Dependabot picks these up too.AGENTS.md(new): install and smoke commands, what the smoke tests do and don't cover, the manual live-validation checklist for dependency-update PRs, and required secret names only.README.md: short "Smoke tests" section that points to AGENTS.md.Live RAG validation stays manual
The smoke tests don't reach a real cluster or the OpenAI API. Per AGENTS.md, dependency-update PRs still need evidence of: app boot plus installed versions, Couchbase and vector index state, that
OPENAI_API_KEYwas available (never its value), the sample PDF and question, the observed RAG, pure-LLM, and cached responses, and screenshots or traces when practical.Local test result
With an invalid key, all 3 provider tests fail, so a bad key can't pass by skipping. CI run 37814472981:
smokepassed 6/6, andOpenAI provider smokepassed 3/3 against the live API.Before the fix,
test_stubbed_app_upload_and_chat[chat_with_pdf_query.py]failed with theEmptyFileErrortraceback rendered in the app. This shows the traceback detection works.cc @nithishr @AayushTyagi1
🤖 Generated with Claude Code