From f96c2ef3485ce5ca6f0d27583e5c11bcc0e3ca0c Mon Sep 17 00:00:00 2001 From: Luiz Henrique Rapatao Date: Mon, 31 Aug 2026 19:51:49 +0100 Subject: [PATCH 1/2] perf(kotlin): resolve operands without compiling a regex `KotlinContext.asValue` built four `Regex` instances inline on the operand path: two in the `when` that decides string literal vs field path, two more in `unwrap`. `process` calls `asValue` on both operands, so a single binary expression compiled up to four patterns on every evaluation. `trim()` also ran twice in the same `when`. The patterns collapse to one hoisted val, the string case is entered once with a single `trim()`, and `unwrap` drops regex entirely. ```kotlin private fun Any?.asValue(): Any? { val result = when { this !is String -> this this == "null" -> null else -> { val trimmed = this.trim() if (QUOTED.matches(trimmed)) trimmed.unwrap() else trimmed.rawValue() } } ... } private fun String.unwrap() = this.trim() .removePrefix("\"") .removeSuffix("\"") ``` `QUOTED` is unanchored. `Regex.matches` is a full-input match, so the `^` and `$` in the original `Regex("^\".*\"$")` were redundant and both `when` branches tested the same predicate. `.` still does not match a newline, so a multiline quoted literal falls through to `rawValue` as before. `unwrap` uses `removePrefix`/`removeSuffix` and not `removeSurrounding("\"")`. Those are not equivalent: `unwrap` strips a leading and a trailing quote independently, `removeSurrounding` strips only when both are present. `rawValue` calls `unwrap` on strings that did not match `QUOTED`, so a key such as `"abc` resolves to `abc` today and would resolve to `"abc` under `removeSurrounding`, changing the map lookup and the throw/no-throw outcome for `OnFailure`. No test changes. --- .../engine/evaluator/kotlin/KotlinContext.kt | 19 +++++++++++++------ 1 file changed, 13 insertions(+), 6 deletions(-) diff --git a/kotlin-evaluator/src/main/kotlin/com/rapatao/projects/ruleset/engine/evaluator/kotlin/KotlinContext.kt b/kotlin-evaluator/src/main/kotlin/com/rapatao/projects/ruleset/engine/evaluator/kotlin/KotlinContext.kt index 57aa495..87ccb26 100644 --- a/kotlin-evaluator/src/main/kotlin/com/rapatao/projects/ruleset/engine/evaluator/kotlin/KotlinContext.kt +++ b/kotlin-evaluator/src/main/kotlin/com/rapatao/projects/ruleset/engine/evaluator/kotlin/KotlinContext.kt @@ -25,10 +25,12 @@ class KotlinContext( private fun Any?.asValue(): Any? { val result = when { - this is String && this == "null" -> null - this is String && !this.trim().matches(Regex("^\".*\"$")) -> rawValue() - this is String && this.trim().matches(Regex("\".*\"")) -> this.unwrap() - else -> this + this !is String -> this + this == "null" -> null + else -> { + val trimmed = this.trim() + if (QUOTED.matches(trimmed)) trimmed.unwrap() else trimmed.rawValue() + } } return when { @@ -64,6 +66,11 @@ class KotlinContext( } private fun String.unwrap() = this.trim() - .replace(Regex("^\""), "") - .replace(Regex("\"$"), "") + .removePrefix("\"") + .removeSuffix("\"") + + private companion object { + // Regex.matches is a full-input match, so no anchors are needed. + private val QUOTED = Regex("\".*\"") + } } From 5e67a6a04397444155e7f3f41d8dbfb92563dfbb Mon Sep 17 00:00:00 2001 From: Luiz Henrique Rapatao Date: Mon, 31 Aug 2026 19:52:29 +0100 Subject: [PATCH 2/2] docs: move benchmark results to BENCHMARKS.md The Performance section was the longest part of the README and is reference material, not something a reader needs while picking an engine. It moves to BENCHMARKS.md with the run-to-run spread of each engine recorded alongside the headline figures, since two of the four configurations vary by more than 30% between runs. The README keeps the engine comparison table and a link. --- BENCHMARKS.md | 106 ++++++++++++++++++++++++++++++++++++++++++++++++++ README.md | 98 ++++------------------------------------------ 2 files changed, 113 insertions(+), 91 deletions(-) create mode 100644 BENCHMARKS.md diff --git a/BENCHMARKS.md b/BENCHMARKS.md new file mode 100644 index 0000000..6effdf9 --- /dev/null +++ b/BENCHMARKS.md @@ -0,0 +1,106 @@ +# Benchmarks + +Measured performance of the engines shipped in this repository, how to reproduce it, and what the numbers mean when +choosing and tuning an engine. See the [README](README.md) for what each engine is and how to use it. + +## Running the benchmark + +Every evaluator module ships a `bench` task that replays the full test rule set against its engine: + +```shell +./gradlew :kotlin-evaluator:bench +./gradlew :rhino-evaluator:bench +./gradlew :graaljs-evaluator:bench -PbenchIterations=5000 +./gradlew :graaljs-evaluator:bench -PbenchIterations=5000 -PbenchReuse=true +``` + +Each iteration evaluates the 147 expressions from `com.rapatao.projects.ruleset.engine.cases.TestData` against the same +input object, after 100 warmup iterations. Results are printed and written to `bench_.txt`. + +Two things to set up before trusting a run: + +* Run at full power. On a laptop in a power saving mode the whole suite lands 25 to 30% low, uniformly across engines. +* Compare only runs from the same session. The harness is a timing loop with no confidence intervals, so a difference + smaller than the run-to-run spread below is not a result. + +## Results + +2000 iterations (294,000 evaluations per engine), Apple M3 Pro, Amazon Corretto 21.0.11. These are relative +magnitudes, not absolute figures: the harness is a simple timing loop, not JMH, and the GraalJS run is interpreter-only +because Corretto is not a GraalVM JDK. + +| engine | ops/s | avg per iteration | p50 | p99 | relative cost | +|----------------------|----------|-------------------|-----------|-----------|---------------| +| Kotlin | 782,637 | 188us | 166us | 331us | 1x | +| Rhino | 352,564 | 417us | 340us | 1.18ms | ~2.2x | +| GraalJS (reused ctx) | ~240,000 | ~590us | ~500us | ~2.0ms | ~3.3x | +| GraalJS | 9,391 | 15.65ms | 15.57ms | 17.36ms | ~83x | + +Each row is one representative run. Run-to-run spread differs sharply by engine, and sets how large a difference has to +be before it means anything: + +| engine | observed across runs | p99 vs p50 | +|----------------------|----------------------|------------| +| Kotlin | 778,000 to 791,000 | ~2x | +| Rhino | 286,000 to 394,000 | ~3.5x | +| GraalJS (reused ctx) | 185,000 to 294,000 | ~4x | +| GraalJS | stable within a few % | ~1.1x | + +The two fastest configurations are the least stable. Once the per-evaluation context cost is gone, an iteration is +short enough that the loop measures JIT and GC noise as much as the engine. Default GraalJS is the opposite: an +iteration is so dominated by context creation that nothing else is visible. + +`GraalJS (reused ctx)` is the same engine with `reuseContextPerThread = true`. Closing the per-call context and +injecting the input into a per-evaluation object costs the default mode about 4% (9,750 to 9,391 ops/s), and buys +deterministic context release plus binding isolation that holds under reuse. + +## Where the time goes + +The `Evaluator` contract sets up a fresh evaluation context on every `evaluate` call. Measuring a single rule +(`item.price equalsTo 10`) separates that fixed cost from the actual rule evaluation: + +| engine | one `evaluate` call | context setup | setup share | +|----------------------|---------------------|---------------|-------------| +| Kotlin | 1.08us | 0.88us | ~82% | +| Rhino | 1.24us | 0.17us | ~14% | +| GraalJS (reused ctx) | 2.19us | 1.40us | ~64% | +| GraalJS | 133.7us | 113.5us | ~85% | + +These rows come from one tight loop over a single rule, after 50,000 warmup calls, so they isolate the steady-state +cost. They are not comparable to the suite numbers above, which include cold and JIT-transient iterations. That loop is +not part of this repository and the `bench` tasks do not reproduce it. The Kotlin row predates the current operand +parsing, which cut per-operand work and not context setup, so its real setup share is above the ~82% shown. + +Reading of the table: + +* **Kotlin**: the fixed cost is flattening the input graph, and it is nearly the whole cost. It scales with the size of + the input object, not with the rule, so a wide input evaluated against a two-field rule pays for every other field +* **Rhino**: setup is entering a `Context`, creating a child scope and injecting the input, because the standard + objects are shared. Before that change the same two columns read 20.5us and 13.0us, a ~63% share. What is left is + compiling and running one small script per operator, which is why deep rule trees cost more than the numbers for a + single rule suggest +* **GraalJS**: context creation dominates almost entirely. On this setup the rule itself is nearly free compared to the + polyglot context it runs in, which is what `reuseContextPerThread = true` removes + +## Practical guidance + +* Reuse the evaluator instance. Operators are resolved once in the constructor, and for GraalJS the shared `Engine` + caches parsed sources across contexts, so a new evaluator per request throws that away +* On GraalJS, set `reuseContextPerThread = true` unless rules are untrusted or deliberately write globals. It is the + single largest win available on that engine +* Pass the narrowest input object that satisfies the rule. All three engines materialise the whole input per call +* Prefer `Map` inputs over arbitrary objects when the data is already in that shape: the object path goes through + Kotlin reflection +* Order `anyMatch` cheaply-first and `allMatch` most-selective-first. Evaluation short-circuits, and with the JS engines + every skipped expression is a script that is never compiled +* On GraalJS, run on a GraalVM JDK (or put the Graal compiler on the runtime classpath) before drawing conclusions from + its numbers. Interpreter-only mode is the default penalty on a stock JDK +* On Rhino, keep the default `interpretedMode = true`. Compiled mode measures about 100x slower, because each operator + compiles a new script that is thrown away + +Both JS engines used to rebuild their whole evaluation environment per `evaluate` call, and that dominated their cost. +Rhino no longer does: it shares one sealed set of standard objects and gives each evaluation a child scope, worth about +7.4x on this suite (3.07ms to 417us per iteration) with no loss of isolation, so there is nothing to opt into. On +GraalJS the equivalent is opt-in because it does trade isolation: `reuseContextPerThread = true` keeps one context per +thread, worth roughly 25x (15.6ms to about 0.6ms), at the cost of rules on one thread sharing a context. Keep it off +for untrusted rules or rules that write globals. diff --git a/README.md b/README.md index c17d72f..6dd849b 100644 --- a/README.md +++ b/README.md @@ -11,13 +11,13 @@ Below are the available engines that can be used to evaluate expressions. All of matter of changing the dependency and the instantiation line. A quick comparison, measured with the benchmark shipped in this repository (details in -[Performance](#performance)): +[BENCHMARKS.md](BENCHMARKS.md)): | engine | operands | throughput (ops/s) | relative | best fit | |-----------|---------------------|--------------------|----------|------------------------------------------------------| -| Kotlin | field paths only | ~566,000 | 1x | high volume, plain comparison rules | -| Rhino | JavaScript | ~350,000 | ~1.6x | rules that need scripting, high volume | -| GraalJS | JavaScript | ~9,700 | ~58x | modern ECMAScript, GraalVM deployments, low volume | +| Kotlin | field paths only | ~783,000 | 1x | high volume, plain comparison rules | +| Rhino | JavaScript | ~350,000 | ~2.2x | rules that need scripting, high volume | +| GraalJS | JavaScript | ~9,700 | ~81x | modern ECMAScript, GraalVM deployments, low volume | ### Kotlin engine implementation @@ -238,93 +238,9 @@ implementation "com.rapatao.ruleset:graaljs-evaluator:$rulesetVersion" ## Performance -### Running the benchmark - -Every evaluator module ships a `bench` task that replays the full test rule set against its engine: - -```shell -./gradlew :kotlin-evaluator:bench -./gradlew :rhino-evaluator:bench -./gradlew :graaljs-evaluator:bench -PbenchIterations=5000 -./gradlew :graaljs-evaluator:bench -PbenchIterations=5000 -PbenchReuse=true -``` - -Each iteration evaluates the 147 expressions from `com.rapatao.projects.ruleset.engine.cases.TestData` against the same -input object, after 100 warmup iterations. Results are printed and written to `bench_.txt`. - -### Results - -Numbers below come from that benchmark, 2000 iterations (294,000 evaluations per engine), on an Apple M3 Pro with -Amazon Corretto 21.0.11. Treat them as relative magnitudes, not absolute figures: the harness is a simple timing loop, -not JMH, and the GraalJS run is interpreter-only because Corretto is not a GraalVM JDK. - -| engine | ops/s | avg per iteration | p50 | p99 | relative cost | -|----------------------|----------|-------------------|-----------|-----------|---------------| -| Kotlin | 566,752 | 259us | 210us | 638us | 1x | -| Rhino | 352,564 | 417us | 340us | 1.18ms | ~1.6x | -| GraalJS (reused ctx) | ~240,000 | ~590us | ~500us | ~2.0ms | ~2x | -| GraalJS | 9,391 | 15.65ms | 15.57ms | 17.36ms | ~60x | - -Most engines are stable under load, with the p99 within 1.2x to 3x of the median. The reused-context GraalJS row is the -exception: it varied between 185,000 and 294,000 ops/s across runs here, so it is quoted as an approximation. Once the -context cost is gone, an iteration is short enough that the timing loop measures JIT and GC noise as much as the -engine. The Rhino row is quoted exactly because it held between 344,000 and 353,000 ops/s across three runs. - -Rhino was measured at 47,827 ops/s (3.07ms per iteration) before it started sharing its standard scope across -evaluations, so that change is worth about 7.4x on this suite. - -`GraalJS (reused ctx)` is the same engine with `reuseContextPerThread = true`. Closing the per-call context and -injecting the input into a per-evaluation object costs the default mode about 4% (9,750 to 9,391 ops/s here), and buys -deterministic context release plus binding isolation that holds under reuse. - -### Where the time goes - -The `Evaluator` contract sets up a fresh evaluation context on every `evaluate` call. Measuring a single rule -(`item.price equalsTo 10`) separates that fixed cost from the actual rule evaluation: - -| engine | one `evaluate` call | context setup | setup share | -|----------------------|---------------------|---------------|-------------| -| Kotlin | 1.08us | 0.88us | ~82% | -| Rhino | 1.24us | 0.17us | ~14% | -| GraalJS (reused ctx) | 2.19us | 1.40us | ~64% | -| GraalJS | 133.7us | 113.5us | ~85% | - -These four rows come from one tight loop over a single rule, after 50,000 warmup calls, so they isolate the steady-state -cost. They are deliberately not comparable to the suite numbers above, which include cold and JIT-transient iterations. - -Reading of the table: - -* **Kotlin**: the fixed cost is flattening the input graph, and it is nearly the whole cost. It scales with the size of - the input object, not with the rule, so a wide input evaluated against a two-field rule pays for every other field -* **Rhino**: setup is now just entering a `Context`, creating a child scope and injecting the input, because the - standard objects are shared. Before that change the same two columns read 20.5us and 13.0us, a ~63% share. What is - left is compiling and running one small script per operator, which is why deep rule trees cost more than the numbers - for a single rule suggest -* **GraalJS**: context creation dominates almost entirely. On this setup the rule itself is nearly free compared to the - polyglot context it runs in, which is what `reuseContextPerThread = true` removes - -### Practical guidance - -* Reuse the evaluator instance. Operators are resolved once in the constructor, and for GraalJS the shared `Engine` - caches parsed sources across contexts, so a new evaluator per request throws that away -* On GraalJS, set `reuseContextPerThread = true` unless rules are untrusted or deliberately write globals. It is the - single largest win available on that engine -* Pass the narrowest input object that satisfies the rule. All three engines materialise the whole input per call -* Prefer `Map` inputs over arbitrary objects when the data is already in that shape: the object path goes through - Kotlin reflection -* Order `anyMatch` cheaply-first and `allMatch` most-selective-first. Evaluation short-circuits, and with the JS engines - every skipped expression is a script that is never compiled -* On GraalJS, run on a GraalVM JDK (or put the Graal compiler on the runtime classpath) before drawing conclusions from - its numbers. Interpreter-only mode is the default penalty on a stock JDK -* On Rhino, keep the default `interpretedMode = true`. Compiled mode measures about 100x slower here, because each - operator compiles a new script that is thrown away - -Both JS engines used to rebuild their whole evaluation environment per `evaluate` call, and that dominated their cost. -Rhino no longer does: it shares one sealed set of standard objects and gives each evaluation a child scope, which runs -the suite about 7.4x faster (3.07ms to 417us) with no loss of isolation, so there is nothing to opt into. On GraalJS the -equivalent is opt-in because it does trade isolation: `reuseContextPerThread = true` keeps one context per thread and -runs the suite roughly 25x faster (15.6ms to about 0.6ms). See [docs/tasks](docs/tasks) for the analysis and the -trade-offs. +The Kotlin engine runs the benchmark suite at about 783,000 ops/s, Rhino at about 350,000, and GraalJS at about 9,700 +on a stock JDK. Full results, how to reproduce them, where the time goes inside each engine, and tuning guidance are in +[BENCHMARKS.md](BENCHMARKS.md). ## Get started