feat(blog): add SQLite FTS5 + Dense hybrid retrieval article - #251
feat(blog): add SQLite FTS5 + Dense hybrid retrieval article#251ishwar170695 wants to merge 3 commits into
Conversation
919d215 to
0e21673
Compare
Signed-off-by: Ishwar <ishwarcm@iitbhilai.ac.in>
0e21673 to
9e914a0
Compare
saiyam1814
left a comment
There was a problem hiding this comment.
Thanks for putting this together. I read through the article and the added assets. Overall, this is a strong technical draft: it has a clear retrieval problem, a concrete architecture, code snippets, benchmark numbers, screenshots, and a useful tradeoff section. I think it can fit the KubeSimplify audience well if we frame it as AI infrastructure / backend retrieval engineering rather than only as a generic RAG article.
A few changes I would suggest before publishing:
-
Use the architecture diagram in the article
The PR addspublic/img/blog/sqlite-fts5-dense-hybrid-retrieval/architecture.png, but the post currently shows the pipeline as an ASCII diagram. This topic benefits a lot from a visual architecture diagram, so I would embed the image in the Architecture Overview section and optionally keep the text explanation around it. -
Fix the author image
public/img/authors/ishwar.jpgappears to be the same file as the blogcover.jpg. That looks accidental. Please replace it with an actual author avatar or a suitable placeholder image. -
Add more context around the benchmark numbers
The reported improvements are compelling (466 msto12 ms,320 MBto48 MB,68%to91%Top-5 retrieval), but they need enough context to be credible and reproducible. Please add details such as corpus size, hardware, Node/SQLite versions, embedding model/vector dimension, how the 100-query benchmark was constructed, and how relevance was judged. -
Verify math rendering for the RRF formula
The post uses$$...$$for the RRF equation. Please verify this renders correctly in the KubeSimplify site. If math rendering is not enabled, convert it to a plain-text/code-block formula so readers do not see raw LaTeX. -
Make the KubeSimplify fit more explicit
I would add a short section on where this belongs operationally: local-first backends, edge apps, small/medium corpora, Kubernetes workloads with persistent storage, serverless constraints, and when teams should move to Qdrant/Milvus/Elasticsearch. This would connect the article more strongly with KubeSimplify’s cloud-native/infrastructure audience. -
Add a note on FTS5 table population/sync
The SQLite FTS5 external-content table example is useful, but readers may wonder howsections_ftsstays in sync withsections. A short ingestion snippet or note about triggers/rebuild flow would make the implementation more complete. -
Polish the style for KubeSimplify
Consider removing emojis from section headings and using a cleaner tutorial/deep-dive style. The current content is good, but simpler headings would make it feel more consistent with existing KubeSimplify technical posts.
Overall: this is close and worth publishing after the above polish. The idea is useful, the implementation is practical, and the performance/tradeoff framing is exactly the kind of detail readers will appreciate once the claims and visuals are tightened up.
|
@saiyam1814 Thanks for the detailed review. I've addressed all the suggested changes:
I also verified that the site builds successfully with these changes. Thanks again for the suggestions! |
|
Did you test this on real hardware? |
Yes. I tested everything on my local machine. |
|
for some reason I am not able to see the preview on cloud flare, let me pull in locally and check |
|
@saiyam1814 @ishwar170695 Here's the preview for the current state of the PR https://post-sqlite-fts5-dense-hybri.website-dab.pages.dev/sqlite-fts5-dense-hybrid-retrieval |
|
@shkatara can you add feedback for this blog? |
|
I'm writing this from my phone so keeping it short. Architecture diagram is not understandable at all to people who are new to this. So many flows and nodes make it confusing than clearer. Flow representation is also not clear. How does one query branch off to different endpoints to search. Where is this configured. FTS5. / BM25. I'm guessing searching something over text so full text search. People not working with db have no idea what this is. Avoid using short forms or at least write full form after them in () Synchronizing the FTS5 Virtual Table: this is worded in a way someone from not a db background would fail to understand. External content table ? What is that. What is a virtual table ? We load only the id (string) and the coordinate list—pre-processed into a compact Float32Array object—into memory: what id is this of ? What is in the float32 array Explain the sql queries in english Because the architecture or packet flow is not clear, I can't put a mental model of why memory usage comes down by 85% What is RRF. What is BM25. The problem statement is good. And one that makes sense. The answer goes into a lot of depth and jargons that makes it hard to follow. I had to try at least three times to read it and always I could not finish. The blog should focus on keeping it simple. Keeping in mind people have no idea what we are writing about and would need help at each layer to put a mental model. If they are fighting to understand what these short forms are, the purpose is defeated. It's better to give them ideas that something like this is possible if the topic is dense, and let them explore their own cases. |
|
@ishwar170695 can you work on the feedback from @shkatara |
… and mental model first
018a736 to
7d7b6d0
Compare
|
@saiyam1814 @shkatara Thanks for the feedback. I've reworked the article to make it more beginner-friendly by simplifying the diagrams, reducing jargon, introducing concepts before acronyms, and focusing on the mental model first. I'd appreciate another review when you have a chance. Sorry for the delay, and thanks for your patience. |
|
Thanks for the rework @ishwar170695, and no worries at all about the delay. The clarity pass really landed. The new pipeline diagram is exactly what @shkatara was asking for: one clean top-to-bottom flow instead of the earlier tangle. Putting the mental model up front before any code was the right call, the I then went through the article side by side with the LawDecoder repo, and a few things have drifted apart between the post and the code. I want to close those before we publish, mainly to protect you: a post like this attracts readers who will clone the repo and run your snippets, and it is much better if everything they find matches. Before we publish1. The FTS5 snippet does not match The post (lines 109-120) shows an external content table: CREATE VIRTUAL TABLE IF NOT EXISTS laws_fts USING fts5(
id UNINDEXED, title, content,
content='laws',
content_rowid='rowid'
);with the caption that Two ways to fix, either is fine:
Also, 2. The repo link promises benchmark scripts Line 228 says the benchmark scripts are on GitHub, but I could not find a harness in the repo, and 3. The reranker deserves its own honest paragraph In const isDocumentForgeryRelated = queryLower.includes('signature') || queryLower.includes('sign') || queryLower.includes('document');
// ... coin/stamp/currency matches: adjustedScore *= 0.01
// ... titles containing forgery/forged: adjustedScore *= 3.0Hardcoded guardrails like this are completely legitimate, plenty of production retrieval systems ship exactly this. The issue is presentation. Right now the article's hook is the one query this rule was written for, so a reader naturally credits FTS5 and RRF for turning counterfeit coins into forgery sections, when part of the credit belongs to the rule. Two options: run a quick ablation over your 100 queries (RRF only vs RRF plus reranker) and publish both numbers, or add a short paragraph saying the reranker is a deliberate domain guardrail for this query class rather than a general component. The ablation would be a great addition if you have the time, since "how much did each stage buy me" is the question every reader will have. One small bug while you are in there: 4. The methodology details from the last round got lost Your earlier reply mentioned the Ryzen 5 5600H, Node and SQLite versions, and the embedding model, but they are not in the current file, I think the rewrite dropped them. Worth putting back, especially the embedding model and dimension ( 5. The memory story needs a breakdown This is the one I got stuck on, and I think it is also what @shkatara meant when he said he could not build a mental model for the 85%. Running the arithmetic: 4,892 sections at 384 dims and 4 bytes each is roughly 7.5 MB for the entire vector cache, and the statutory text for 4,892 sections is maybe 10 to 20 MB. Neither of those explains a 270 MB drop. My guess is the heap is actually dominated by the transformers models ( If that is right, the fix is easy and the post gets more interesting, not less: show a rough heap breakdown, and say the win came from dropping a full JSON parse and keeping typed arrays instead of object graphs. If I have the wrong end of it, a couple of numbers in the section will settle it. 6. Name the real cause of the latency win The table labels v1 as a linear JSON scan, which is honest, but the surrounding narrative lets the reader attribute 466 ms to 12 ms to hybrid retrieval. Most of that gap is really "we stopped re-parsing a JSON file on every query" and would have shown up even without FTS5 or RRF. One sentence saying so keeps the claim solid, and the hybrid architecture still has plenty to stand on with the accuracy result. Polish
Where that leaves usThe 6 items in the first section are what I would like fixed before we merge, and 1, 2, and 3 are the important ones since they are what a reader would notice when they open the repo. The polish list is quick and mostly mechanical. To be clear about the overall read: the idea is good, the engineering is real, and the writing is a big step up from the first draft. This is worth publishing once the post and the code tell the same story. Ping me when you have pushed and I will do another pass. |
Changes proposed
Adds a new, technical, practitioner-led article to the Kubesimplify blog focusing on hybrid retrieval architectures in local RAG systems.
content/blog/sqlite-fts5-dense-hybrid-retrieval.md— Deep-dive into sparse vs. dense search limitations, SQLite FTS5 BM25 configurations, Reciprocal Rank Fusion (RRF), deterministic domain reranking, and vector cache memory footprints.public/img/blog/sqlite-fts5-dense-hybrid-retrieval/— Static assets (architecture diagrams, benchmarks, UI screenshots).content/authors.json— Add author profile entry forishwar.public/img/authors/ishwar.jpg— Author avatar placeholder.No other changes are made to the site's code.
Note to reviewers
This post focuses on systems-level search engineering and database schemas (SQLite), which is well-suited for KubeSimplify's backend, cloud-native, and infrastructure audience.