Skip to content

Bugs found by batch 2 of the test programs, on every engine - #17

Merged
YoanSallami merged 19 commits into
feat/v2-assertionsfrom
test/more-programs-2
Oct 7, 2026
Merged

YoanSallami merged 19 commits into
feat/v2-assertionsfrom
test/more-programs-2

Conversation

@YoanSallami

Copy link
Copy Markdown
Contributor

This adds batch 2 of the self-checking programs: 136 of them, plus 13 for the bugs they found. They cover functions, functors, conditionals, booleans, arithmetic, aggregates, ordering, search, nulls, strings, arrays, joins, negation, recursion, records and unions. Each is mirrored into the compiler and verifier fixtures and run on every engine.

The full e2e suite passed locally on all six engines (8,374 tests):

  • 13 tests failed in the full run. Six were the Spark unnest bug, which is now fixed. Seven were Presto aborting queries when the Mac slept and the VM clock jumped.
  • All 13 pass when rerun, and all Databricks tests pass after the unnest fix.

Fixed

Semantics

  • / divides exactly everywhere. 7 / 2 was 3 on SQLite, PostgreSQL, Trino and Presto.
  • A minus after an operator now parses: 2 * -3, 7 % -3, 2 ^ -1.
  • @OrderBy and @Limit apply to a recursion's result, not to each step. A limit used to cut every step, and a -1 recursion could loop forever. Semi-naive recursions also lost their ordering.
  • NULLs sort last on every engine.

Checks

  • A function written as a condition (V(x:), IsEven(x)) used to hold for every row; the verifier now refuses it.
  • A functor's result can be read by later rules; it was reported as undefined.

Engine functions

  • Format renders %d, %.Nf, widths and %% on Presto, PostgreSQL and Trino.
  • Join uses ARRAY_JOIN on Trino, Presto and Databricks.
  • ToString of a double is no longer scientific on Trino.
  • Range(0) is empty everywhere.
  • Size([]) and ArrayConcat(l, []) no longer need an array type.
  • ToInt64 on SQLite no longer repeats its argument, which made nested conversions overflow SQLite's parser.

Paging and search

  • Results are ordered again outside the subquery; Trino and Presto dropped the order.
  • Trino puts OFFSET before LIMIT.
  • Presto, which has OFFSET disabled, numbers the rows instead.

Correlation

  • A negation inside a negation's subquery now reads its parent's table. Trino, Presto and Spark can't correlate two levels up.
  • A table alias never equals a column name, ignoring case. Trino read K.k as the column k.

Databricks

  • Recursions are computed into tables, because Spark's planner is exponential in the unrolled steps.
  • Unnests use a LATERAL subquery.

Other

  • introspect leaves out Synalog's own working schemas.
  • The e2e Spark connection times out instead of hanging the run.

Each change in the SQL is documented in tests/compiler_tests/DEVIATIONS.md. The compiler harness compares goldens without NULLS LAST, and a unit test checks that each dialect emits it.

🤖 Generated with Claude Code

synalinks-saas and others added 19 commits October 5, 2026 05:55
136 programs (functions, functors, conditionals, booleans, arithmetic,
aggregates, ordering, search, nulls, strings, arrays, joins, negation,
recursion, records, unions), mirrored into the fixtures and run on every
engine. Fixed:

- `/` divides exactly everywhere (7 / 2 was 3 on SQLite, PostgreSQL,
  Trino and Presto).
- A minus after an operator parses (`2 * -3`, `7 % -3`, `2 ^ -1`).
- A function written as a condition is refused (it held for every row).
- A functor's result is defined for the rules that read it.
- @orderby and @limit apply to a recursion's result, not to each step
  (a limit cut every step; until convergence it could loop forever); the
  semi-naive result keeps its ordering.
- Format renders %d, %.Nf, widths and %% on Presto, PostgreSQL and Trino.
- Pages and searches order again outside their subquery (Trino and Presto
  dropped the order); OFFSET before LIMIT on Trino; numbered rows on
  Presto, where OFFSET is disabled.
- Nulls sort last on every engine.
- Join is ARRAY_JOIN on Trino, Presto and Databricks; ToString of a
  double is not scientific on Trino; Range(0) is empty; Size([]) and
  ArrayConcat with [] need no array type; ToInt64 on SQLite names its
  argument once (nested conversions overflowed SQLite's parser).
- A negation inside a negation's subquery reads its parent table
  (Trino, Presto and Spark cannot correlate two levels up); a table
  alias never equals a column name (Trino read `K.k` as a column).
- Databricks computes recursions into tables (Spark's planner is
  exponential in unrolled steps) and unnests in a LATERAL subquery.
- introspect leaves out synalog's own schemas.
- e2e: the Spark connection times out instead of hanging the run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
104 programs (combine expressions, ordered and grouped aggregates, strings,
nulls, booleans, casts, math, arrays, records, joins, negation, recursion,
functors, directives, dates, unions, functions), mirrored into the
fixtures. Fixed:

- ToString of a boolean is "true" or "false" on SQLite too (it wrote 1
  and 0).
- PostgreSQL: a combine whose operand type is unknown is not cast (Min=
  over strings was cast to numeric); a null of a column of known type is
  cast to it (null UNION null is text there, and failed against numbers).
- Column types no longer depend on hash order (a rule giving a column a
  null made it unknown at random).
- A run drops the tables it computed in synalog's own schemas (they
  filled Presto's memory connector until its heap ran out).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
About 300 programs (empty predicates, nested aggregates, mixed numbers,
generated grids of operators, comparisons, aggregates, string functions,
nulls, recursion shapes, memberships, pages, negations, joins, casts,
formats, records, functors, booleans), mirrored into the fixtures. Fixed:

- A function the program defines is the one a value calls, also when a
  built-in has its name (`Size(x) = ...` compiled to the built-in).
- ToInt64 of text on SQLite is a plain CAST: SQLite before 3.46 (CI's
  3.37) overflowed its parser stack on nested conversions.
- Databricks writes constant facts as one VALUES (Spark cannot plan a
  correlated subquery over a union of constant SELECTs); records stay a
  union.
- Split("") is one empty part on PostgreSQL; Size of a null list is null
  on PostgreSQL (CARDINALITY) and Databricks (ARRAY_SIZE); Split of a null
  is null on SQLite.
- e2e: the Presto container keeps finished queries a minute, with a 2 GB
  heap (the suite's queries filled its 1 GB heap until it exited).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rtions

- Batch 6: generated grids (recursion shapes, ArgMin/ArgMax, Array=,
  StringAgg=, combine, Format, records, functors, unions, searches, pages,
  booleans, Like patterns, precedence).
- Batch 7: 101 programs of the knowledge-graph patterns of
  docs/knowledge-graphs.md (nodes, edges through nodes, inverse, symmetric,
  typed, composed, weighted and reified edges, traversals, neighborhoods,
  restriction by functor, cardinality and cycle checks, temporal graphs),
  and 30 assertion programs over the same graph (integrity, keys, inverse
  and symmetric edges, closures, acyclicity, temporal intervals).

Fixed:
- The parser panicked on a character of several bytes (a program starting
  with `é`, a statement string with `∀` and a quote): a span edge inside a
  character is moved to its boundary. Fuzzing found no panic since.
- Trino computes recursions into tables: it copies an unrolled recursion
  into every read, past its limit of 150 stages when a query reads it
  several times (an assertion reading a closure three times).
- Split("") is one empty part on PostgreSQL; Size of a null is null on
  PostgreSQL and Databricks; Split of a null is null on SQLite.
- Databricks keeps facts holding records as a union (Spark refuses a
  record in a VALUES row).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…/Trino; batch 8

- Verifier: a rule whose comparisons can never all hold (a < b, a >= b;
  x == 1, x == 2; a 3-cycle of <) is reported as a warning, since it gives no
  row. A rule with a disjunction is reported only when every branch
  contradicts itself. Documented in docs/verification.md.
- Presto/Trino: a derived predicate read more than once is materialized, as
  @Ground does, so a knowledge graph's neighborhoods stay under Presto's
  stage limit.
- e2e: Trino's memory connector hung on a join with an empty side (lazy
  dynamic filtering); turned off in the test compose file.
- tests/programs: bill of materials, fraud rings, approval chains, their
  assertions, and contradictions (batch 8), mirrored into the compiler
  fixtures; programs can state `# Expect: quiet` (no warning).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Compiler and parser:
- A condition comparing a variable with an expression of itself
  (`last == Upper(last)`) was dropped as "self-referential".
- Record fields take their value's type (a column's, a list's element, a
  number for arithmetic...): Trino, Presto and PostgreSQL typed every
  non-literal field as text, so `{n: x}` sorted "10" before "9".
- A head value after `=` may compare (`F(x) = if x >= 4 then ...`,
  `n? += if x != 0 ...`); `x + 1 == y` in a body stays a comparison.
- A parenthesized combine is a positional argument
  (`Coalesce((combine += 1 :- P(a: n)), 0)`).
- A grounded predicate named after a keyword (`Order`, `In`) gets the table
  `<Name>_table`; a column `at` is quoted (DuckDB's AT TIME ZONE).
- Databricks `Split` takes its separator as text, not a regular expression.
- `Round` and `ToInt64` round a half away from zero on DuckDB and PostgreSQL
  too; PostgreSQL's `Round(x, digits)` of a double works.
- A functor applied to a recursion copies its loop on Presto, Trino and
  Databricks (it ran the original's loop over the original's edges).

Verifier and assertions:
- A combine compared (`x > (combine Avg= w :- P(w:))`) binds its own
  variables: no false "unbound variable".
- Statements: a nested formula may compare an outer variable
  (`∃ u, P u ∧ u ≠ x`), a ∀ may nest in an ∃, a variable a function is applied
  to ranges over where the function is defined, `(P x y)` keeps its
  arguments. Docs say which variables name the counterexamples' columns.

Runners: the PostgreSQL `ARRAY_CONCAT_AGG` is created once, safely when two
sessions start together. Tests: Trino's heap capped at 3 GB; the golden
generator compiles the engines in parallel.

Programs: batches 9-19 (inventory, org chart, series, contradictions, social,
ledger, text, records, quotes, numbers, lists, aggregation, functions, nulls,
pages, library, shipping, functors over recursion, dates, clock, joins, order,
three batches of @Assert statements, graphs), mirrored into the compiler
fixtures with their goldens and deviations documented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…es 18-19

- Today and Now read the clock in UTC on every engine, whatever the
  session's time zone (DuckDB, a Trino client and Databricks used local
  time, so the hour and, near midnight, the date differed by engine).
- Trino's ToString of a timestamp is `2026-10-05 15:19:59.910`, as on the
  other engines, not ISO's `T` form.
- A parenthesized combine is a positional argument
  (`Coalesce((combine += 1 :- U(a: n)), 0)`): the colon of `:-` was read as
  a field's.
- tests/programs: a third batch of @Assert statements (roads: functions of
  two arguments, primes and Greek names, float sums, keyword columns,
  recursion and functors) and undirected graphs, with their fixtures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- tests/programs/modules: imports with and without aliases, imported
  predicates in joins, negation and aggregation, generic aggregates and a
  recursion applied across an import by functors, order, search and
  assertions over imports, import errors; its library in modules/lib.
- tests/programs/pipelines: six layers of predicates, @Ground on intermediate
  ones, a predicate read up to four times, unions and negations stacked,
  chains of up to 18 predicates.
Both mirrored into the compiler fixtures with their goldens.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion reports

SQL injection (tests/programs/injection, 180 attempts on every engine):
- Databricks and BigQuery string literals write a quote of the value as
  ", so a literal holds no quote but its own two and no statement
  splitter ends it early ("\"; DROP TABLE t; --" ran DROP TABLE t).
- search() patterns are the dialect's string literals (a backslash ended
  them on Spark and BigQuery).
- PostgreSQL strings with a backslash are E'...' literals.
- @dataset takes a schema name (verifier and compiler); @AttachDatabase and
  copy_to_file paths are escaped; a table named in @Ground lives in
  Synalog's dataset, so no program drops a table elsewhere.

A number's text: ToString of a number is the same on every engine: a whole
number without a decimal point, every digit below 10^18; otherwise at most
15 significant digits in plain decimal, trailing zeros trimmed (0.1 + 0.2 is
"0.3", 1e20 is "100000000000000000000"). The value is named once, so nested
conversions stay small. Type inference knows the result types of arithmetic,
aggregates and string functions.

Assertion reports: counterexample values from the database are shown as
quoted, escaped, bounded literals and statements on one line
(synalog.quote_value, synalog.statement_text): data read by an agent cannot
pass for report text.

Also: Like escapes with a backslash on every engine; the verifier refuses a
disjunction inside a negation or a combine; batches 22-24 (strings,
analytics, injection) with their fixtures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Batches 25-38 (about 1,400 programs in tests/programs, mirrored as compiler
fixtures with goldens) and the bugs they found, fixed in the compiler.

Security
- A program reaches no file, process or service: the library's ReadFile,
  ReadJson, WriteFile, PrintToConsole, Intelligence and Clingo functions are
  gone, @AttachDatabase and @Ground(copy_to_file:) are refused, and SQLite
  sessions refuse Logica's file functions.
- A table name is a name: `(SELECT ...)` and any other text in backticks
  was written into the SQL as such; a quoted path is quoted part by part.
- `${flag}` is not substituted into the SQL text (it was, inside string
  literals too); FlagValue() gives a flag as a literal.
- Spark substitutes `${var}` in a statement's text: a dollar in a Databricks
  literal is $, and a field or predicate name holds no `$`.
- Databricks quotes a backtick by doubling it, BigQuery escapes a backslash.

Language and verifier
- A bare boolean variable or field is a condition (`..., active`).
- A variable only tested (`!flag`, `flag && x`, `Like(s, ...)`) is unbound.
- A variable a negation binds and uses again is the negation's own.
- A function named like a built-in one is refused; aggregates called outside
  an aggregation, and ArgMaxK / ArgMinK without a count, are refused.
- A combine's variables in a head are its own; a predicate a head's combine
  reads is checked as defined; FlagValue of a flag @DefineFlag does not
  define is refused.

Semantics, the same on every engine
- Numbers with a point are doubles (DuckDB and PostgreSQL read decimals:
  `0.1 + 0.2 == 0.3`, DECIMAL overflows); Round(x, d) rounds the number at 15
  significant digits, half away from zero; a division or remainder by zero
  is null; SQLite's `%` keeps fractions, PostgreSQL's takes doubles.
- A null meets no null: two null keys no longer join.
- Join on SQLite in SQL (a null element is skipped); Repeat, StartsWith,
  EndsWith of a null on Trino, Presto and SQLite; RegexpContains of a null.
- search() reads a column in the text ToString gives it.
- Div tells the quotient's sign without a product that overflows.
- A large value of a variable used more than once is computed once: chained
  definitions no longer grow the SQL exponentially (2 MB for a date function).

Types
- FlagValue is text.
- Type inference shares a vertex's type across the edges it is in: a
  relation's column wherever it is read, a rule's variable, a field of
  either; a column's type is the intersection of its rules', in any order.
- Empty and null lists are typed from where they are written (PostgreSQL,
  Trino, Presto, Databricks); a record takes its column's field types.
- Lists of records: composite arrays on PostgreSQL, array(row) on Trino and
  Presto, UNNEST keeping a record one column; a composite type's name hashes
  its declaration; Trino's and PostgreSQL's reserved words are quoted.
- A recursion's accumulated table is created with the widest types its steps
  can give (steps applied to empty tables), so every step inserts its new
  rows, in linear time, on every engine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… they missed

Batches 39 to 51 (about 1300 programs) over negation, conditionals,
unions, combines, grouped aggregation, joins, text, numbers, lists,
disjunctions, user functions, records and ordering, on every engine. The
last four batches found nothing.

What they found, harmonized on every engine:
- Min= and Max= of booleans on PostgreSQL (BOOL_AND, BOOL_OR).
- Greatest and Least are null when an argument is.
- Element is null below zero and past the end (Spark's index cast to INT).
- List=, Set=, Array= and ArgMaxK over no rows are null; Spark collects a
  null value as the other engines do.
- Reverse on SQLite, written in SQL.
- A combine over the elements of the outer row's lists compiles to array
  lambdas on Trino, Presto and Spark, which refuse that correlation.
- The outer rule's variable types reach a combine; a collected list has the
  type of its elements; a type-inference prefix collision.
- Facts with a typed null stay one VALUES on Databricks; Range casts to an
  integer there.
- The parser kept the parentheses of a lone combine argument; the verifier
  no longer flags a combine inside an aggregate.

Docs: Count=, nulls and empty groups, combine, lists and `in`, record
fields, disjunction rows, functions with a body, Ltrim, Rtrim, Reverse,
Pow, Greatest, Least, Element's range, default null ordering and ties.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The e2e step ran every engine in one job and outgrew its 45 minutes: each
run was cut off with every test passing so far. Each engine now runs in a
job of its own, in parallel, starting only its engine's server, with three
hours for its tests. SYNALOG_E2E_ENGINES (a comma list) selects the engines
whose tests run; the introspection tests without an engine parameter say
their engine with a marker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A null literal keeps its own type, any type: giving it another changes
nothing, yet it was counted as a change, so the inference loop never
converged and ran its 1000 iterations whenever a null met a typed value (a
column of numbers with a null among them). Such a program compiled twelve
times slower (440 ms instead of 36 for eight facts), which made most of
the end-to-end suite's time. A change is now the vertex's type changing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Measured on the 4910 golden fixtures, in release builds: 44.2 s in all
(7.7 ms median) before, 8.4 s (1.3 ms median) after. The SQL is the same:
every golden matches on every engine.

- The built-in function tables (bulk, base, per dialect), the basis
  functions and the built-in names were rebuilt for every rule translated:
  they are built once per dialect and shared.
- The dialect's library program was parsed for every program: once per
  dialect.
- Type inference looked a vertex's type up by formatting the vertex as text
  and hashing it, in every step of its loop, and copied each edge in every
  step: vertices are interned once, their types kept in a vector.
- The column types of the predicates were found by a pass over all the
  edges per predicate: one pass for all of them.
- The type graph kept every edge twice in maps of maps keyed by text, and
  rebuilt its list of edges from them: a list of its edges, in the order
  they were connected, and a set of their keys. A merge of the graphs that
  added nothing is removed.
- Recursion unfolding rebuilt its index of all the rules after each
  predicate it made: it adds the new rules.
- Eliminating a variable copied the keys of every node it walked, and
  counted the variable's uses for every replacement: it walks the values,
  and counts only for a large value.
- Building a program cloned the rules and their heads several times: it
  borrows them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
7.6 s for the 4910 golden fixtures (from 8.4 s), the slowest ones 15 to
25 % faster; every golden still matches on every engine.

- The functors kept the program's rules twice and cloned their index:
  once, and the index is moved.
- The type graph formatted both expressions of every edge as text to know
  it once: a 128-bit digest of their text, written without allocating it.
  It keeps the columns its edges hold, the one thing asked of its
  expressions.
- Variable elimination copied both sides of every unification on every
  pass: it reads them in place and copies the value it assigns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI's SQLite (3.37, Ubuntu 22.04) failed five programs with "parser stack
overflow": it parses about 30 nested calls at most, and the text of a
number took twelve, often inside other calls (`Array= id -> ToString(amt)`).
The text takes three fewer, the same on 5032 values: the guards its last
branch never needs are gone, and the exponent of the large numbers is read
at its place. A search test expected a number cast as text where it is
read as ToString writes it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ToString of a number (and the text search reads) is the number as written,
its shortest text that reads back as the same double, rounded half away from
zero to 15 significant digits but at most 15 decimals, as Round(x, digits)
reads it. Each engine rounded the double through its own printf, format or
ROUND, and they differed: 342547.0843250365 was ...036 on DuckDB and
PostgreSQL and ...037 elsewhere; a tie went to even on two engines; from
10^14 to 10^15 three kept 16 digits; floor(log10(x)), rounded up just below
a power of ten, made Trino write 999999999999999.4 as 1000000000000000;
PostgreSQL's numeric of a double kept 15 digits of 9999999999999998; SQLite
3.37's printf drifted in the last digits.

The shortest text is read as a DECIMAL, its integer digits counted, and it
is rounded exactly. DuckDB and Spark take a constant number of places in
ROUND: one per count of digits. SQLite has no exact decimals: its sessions
register SYNALOG_NUMBER_TEXT, which follows the rule. All six engines give
the rule's text on 992 values; tests/programs/numtext holds the cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The CLI tests failed on macOS ARM with Python 3.10: its SQLite (before
3.43, without an extended long double) reads a float literal by rounding its
digits to a double and scaling by a power of ten, so 4503599627370495.5 read
4503599627370495. A literal whose digits pass 53 bits, or whose power of ten
passes 22, is written as its double's exact form, an integer of at most 53
bits scaled by powers of two; other literals are written as they are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A semi-naive step wrote its new rows into one table, then copied them into
the table the next step reads, and counted them: six statements a step on
Trino, Presto and Spark, where each statement costs a query's fixed
overhead. The steps now alternate the two tables, each reading one and
writing the other, and convergence is counted once every two steps: three
statements and a half a step. An odd number of steps ends with one more
step after the loop. The 15 slowest recursion tests run 24% faster on
Presto and 17% on Spark; every recursion folder gives the same rows on all
six engines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@YoanSallami
YoanSallami merged commit 45be0a3 into feat/v2-assertions Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants