Skip to content

Prototype smart search for Indonesian, with word lists inside the Bible version - #298

Open
yukuku wants to merge 7 commits into
developfrom
claude/search-tips-communication-lj6181
Open

yukuku wants to merge 7 commits into
developfrom
claude/search-tips-communication-lj6181

Conversation

@yukuku

@yukuku yukuku commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Why

Search matches letters, and for Indonesian that fails in both directions:

  • It misses. kasih can't find mengasihi, because the prefix swallows the k.
  • It over-matches. berkat returns every berkata.

This prototype widens each plain query word to its whole word family and matches whole words. On Terjemahan Baru, in verses:

Query Letter search Smart search
sembuhkan sakit 6 48
beriman 14 225
berkat 3,059 279

Word lists live in the Bible version

The word families are part of the version's own data, like xrefs and footnotes.

Where they're stored

  • In a .yes file, as a new lexicon section.
  • In the internal version, as {prefix}_lexicon_bt.bt.
  • Both use the same Bintex format and are read by LexiconSection (AlkitabYes2).
  • Version.loadLexicon() exposes them, implemented by Yes2Reader and InternalReader.

How forms are encoded (data format version 1)

  • The section holds start rules, end rules, a pieces table, then the families. Each form is a sequence of tokens, with no separator characters:

    Token Meaning
    0 the root as is
    1 the root with its start rewritten by the start rules
    2 a literal, followed by an autostring
    3 the root with its end rewritten by the end rules
    4, 5 reserved, rejected by the reader
    n ≥ 6 entry n − 6 of the pieces table
  • Any text can sit between roots: mereka-rekakan is me, root, -, root, kan.

  • End rules cover English spelling changes (love → loving, carry → carried) the same way start rules cover Indonesian nasal prefixes.

  • For TB (2,942 families, 15,373 forms) that is 78 KiB as the internal file, or 49 KB inside the .yes after Snappy.

  • The spec is in docs/binary-formats.md.

.yet and the converters

  • .yet gains lexicon lines: the root, then every form spelled out, tab-separated. No notation, no rule table.
  • LexiconCompiler, called by the section writer, decides the encoding. It adopts rewrite rules one at a time, each time the one that makes the output smallest, then splits every form into root tokens and text. Text runs used at least twice become shared pieces, and the rest are literals.
  • For TB it picks t → "", s → y, k → g, p → m with men as a shared piece, not the textbook k → ng. The file is the same size as with the hand-written table.
  • YetFileInput, YetToYes2 and YetToInternal carry the families over and print what the compiler chose. CopyProprietaryAssetsTask now packages *_lexicon_bt.bt.

Related PRs

  • TB's internal file: yukuku/androidbible-proprietary#1
  • The lexicon lines in in-tb.yet: yukuku/alkitab-sources#3

An Indonesian version without a lexicon gets rules-only families derived on the device.

How a query term is resolved

Checked in this order:

  1. The version's word list.
  2. The version's vocabulary: a word with no relatives matches only itself.
  3. Affix peeling, for words the text never uses, e.g. sembuhkan → sembuh.
  4. The classic letter match.

+word and quoted phrases are unchanged.

Diagnostics (visible to the user, on by default in the prototype)

Panel above the results

  • how each word was resolved, and the peel trail;
  • every form searched, with counts;
  • forms the letters can never reach are highlighted, and forms absent from this version are faded;
  • timings;
  • gained/dropped verses compared with the letter search, with All / New / Dropped filters.

Result rows

  • NEW / DROPPED badges, plus a via … line naming the forms that matched.

Search Lab (search menu)

  • the switches and the word-list mode: automatic / built-in only / rules only;
  • the selected list and vocabulary for any version;
  • a "try a word" box comparing the built-in list with rules only;
  • a self-check over known-tricky queries.

Also: two experimental settings (Smart search (Indonesian), Smart search diagnostics), and rewritten search tips (EN/ID).

Testing

  • ./gradlew testPlainDebugUnitTest testPlainReleaseUnitTest :AlkitabYes2:testDebugUnitTest assemblePlainDebug: all 1,738 tests passing.
    • 855 debug unit tests.
    • 28 AlkitabYes2 tests, including 11 in LexiconSectionTest: rule matching at both ends, the rules found for Indonesian and English samples, text between roots, pieces vs literals, reserved tokens, round trips.
  • TerjemahanBaruSmartSearchTest checks the table above against the real TB .yet. It also checks that the generated in-tb.yes section and tb_lexicon_bt.bt decode to the same families. It needs ALKITAB_TB_YET (plus optionally ALKITAB_TB_YES / ALKITAB_PROPRIETARY_DIR) and is skipped otherwise; locally it passes 9/9.
  • The converters were compiled and run locally on in-tb.yet. Every regenerated tb_* internal file is byte-identical to the one in the proprietary overlay, apart from the new lexicon file.
  • SmartSearchSnapshotTest renders the panel and Lab in EN/ID to Alkitab/build/snapshots/smart-search/.
  • Not run on a device or emulator. The sandbox has no KVM.

Open question

The TB word list keeps kasihan as its own family, so kasih drops 145 verses, mostly belas kasihan. The Dropped filter shows them.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Xk4Y839XNeb9v1wkQDpQWm

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

📱 Preview builds

Signed release builds of 5ec5ad0 — version 5.0.0-dev.121 (23848990) — from this run.

Flavor Application ID APK
Alkitab yuku.alkitab Download (7.7 MB)
Quick Bible yuku.alkitab.kjv Download (7.5 MB)
Sabda Alkitab org.sabda.alkitab Download (7.7 MB)

Or open https://182910cb-alkitab-pr.yukuku.workers.dev on an Android device.

These share their application IDs and signature with the Play Store builds, so installing one replaces the corresponding installed app (data is kept).

This comment tracks the latest build for this PR; earlier builds keep their own URLs.

…ble version

Letter search fails Indonesian both ways: kasih cannot find mengasihi
because the prefix swallows the k, and berkat returns every berkata.
Smart search widens each plain query word to its word family and
matches whole words. A word the version never uses has its affixes
peeled until a listed form is reached, and anything unknown still falls
back to the letter search. +word and quoted phrases keep their exact
meaning.

The word families are part of the version's own data, like xrefs and
footnotes: a lexicon section in yes files and <prefix>_lexicon_bt.bt for
the internal version, one Bintex format read by LexiconSection. Forms
are abbreviated: ~ stands for the root and < for the root rewritten by a
prefix table stored alongside (k -> ng), so kasih me<i is mengasihi.
The .yet format gains lexicon_prefix and lexicon lines, which YetToYes2
and YetToInternal carry over. An Indonesian version without a lexicon
gets rules-only families derived on the device.

Diagnostics are visible in the app while the prototype is evaluated: a
panel above the results explains how each word was resolved and which
forms were searched, compares with the letter search (new and dropped
verses, with filters), and reports timings; result rows are badged NEW
or DROPPED with the forms that matched. The Search Lab holds the
switches and word-list mode, plans any query with both the built-in
lexicon and rules only, and runs a self-check. The search tips describe
the syntax actually in effect.

On Terjemahan Baru: sembuhkan sakit goes from 6 to 48 verses, beriman
from 14 to 225, and berkat from 3,059 to 279.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xk4Y839XNeb9v1wkQDpQWm
@yukuku
yukuku force-pushed the claude/search-tips-communication-lj6181 branch from a6e56fa to f164385 Compare September 23, 2026 03:59
@yukuku yukuku changed the title Prototype smart search for Indonesian, with in-app diagnostics Prototype smart search for Indonesian, with word lists inside the Bible version Sep 23, 2026
claude and others added 6 commits September 23, 2026 05:58
Each form in the lexicon section is now a sequence of tokens: 0 is the
root as is, 1 is the root rewritten by the prefix table, 2 is a literal
followed by an autostring, 3 to 5 are reserved, and n >= 6 is entry
n - 6 of a table of text pieces. Only pieces used at least twice go in
the table; the rest are literals. No separator characters are needed,
and any text can sit between roots (me, root, "-", root, kan).

The .yet keeps the readable ~ and < notation, now tab-separated, and
the converters encode it into tokens.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xk4Y839XNeb9v1wkQDpQWm
…them

A .yet lexicon line is now the root and its forms spelled out, with no
notation and no prefix table. LexiconCompiler, used by the section
writer, finds rewrite rules for both ends of roots in the data, adopting
one at a time whichever makes the output smallest, then splits each form
into root tokens and text and shares the text runs used more than once.

The section format is version 1 (nothing has shipped): start rules, end
rules, pieces, then families. Token 3 is the root with its end rewritten
by the end rules, so English changes such as love to loving or carry to
carried fit the same scheme as Indonesian nasal prefixes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xk4Y839XNeb9v1wkQDpQWm
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants