Existing import tt code must keep working. Below is everything that changed,
and why.
tt.vector / tt.matrix and their attributes d, n, r, core, ps, erank, is_complex, from_list / to_list / full / round / norm / copy, the arithmetic,
__getitem__, and the whole zoo of constructors (ones, rand, eye, xfun, linspace, sin, cos, delta, stepfun, unit, qshift, IpaS, Toeplitz, qlaplace_dd, kron, mkron, zkron, zkronv, zmeshgrid, zaffine, concatenate, sum, reshape, permute, matvec, col, diag, dot).
The index order is preserved: mode 1 is the fastest (flat = i1 + n1*i2 + ...),
a TT-matrix core is (r, i, j, r) with the row index first, and the merge is
s = i + n*j. x.core and x.ps give exactly the same F-ordered layout as
before (tests/test_core.py::test_core_and_ps_match_legacy_layout checks it
byte for byte).
The owner of the truth is now x.cores (a list of (r,n,r) arrays). x.core
and x.ps are computed properties. Assigning x.core = buf rebuilds the
cores, but requires n and r to be known already; the old idiom — create an
empty tt.vector(), set d, n, r, call get_ps(), drop in core — no longer
works. Use tt.vector.from_list(cores) or tt.vector.from_flat(core, n, r)
instead. The reason: two owners of the same numbers drift apart sooner or later.
Toeplitz,IpaSandqshiftreturned the transposed matrix.tt.qshift(d)now gives the lower shift (ones on the first subdiagonal),tt.IpaS(d, a)the lower bidiagonal withabelow the diagonal, andToeplitz(x, kind='L')the lower triangular one. Verified against dense references. If your code compensated for the old behaviour with a transpose, remove the compensation.IpaSandqshiftwere also rebuilt through an explicit carry construction (binary addition) rather than throughToeplitz.reshapeof a TT matrix now cuts rows and columns in step. The flat reshape mixed row bits with column bits and produced a silently wrong result.tt.dot(a, b)conjugates its first argument (sum(conj(a) * b)). Nothing changes for real data.
Local systems smaller than this are solved densely, larger ones by GMRES. The
value 50 made sense where the local solver was Fortran. In Python the threshold
is a different number: on a max_full_size=50 explicitly.
x.full()for a block TT (r[0] > 1orr[-1] > 1) returns the shape(r0,) + n + (rd,)with unit boundaries dropped. The old version claimed the shape[rd, r0, n...], which did not match the layout.tt.tensoris a deprecated alias oftt.vectorand warns viaDeprecationWarning.write()/read()write.npzinstead of the binary tt-fort format.rand(n, d, r, samplefunc=...):samplefunc(size)is called as before, but the default generator is numpy'sdefault_rngrather than the globalnp.random.six,np.floatandnp.complexare gone — the package runs on numpy >= 1.24.- Sample-DIRT's reusable kernel is available from
tt.transportand lazytt.SampleDIRT-style imports remain supported. Its research experiments, reports, and generated artifacts live in the separatesample-dirtproject, which imports ttypy rather than carrying a second implementation.
The old package vendored two foreign projects. Both were replaced and checked against dense truth:
| was | size | replaced by | verified against |
|---|---|---|---|
EXPOKIT (tt-fort/expm/, dexp_mv, normest) — the local matrix exponential in KSL |
5803 lines of F77/F90 | expmv_krylov + _norm_estimate in tt/algs/ksl.py (Arnoldi with adaptive substepping), ~50 lines |
scipy.linalg.expm |
PRIMME (tt-fort/primme/) — the local eigenproblem in eigb |
89 C/Fortran files, ~1 MB | a dense eigh for small blocks, scipy.sparse.linalg.lobpcg for large ones |
the analytic eigenvalues of the Laplacian |
Measured on one machine, with bit-identical input.
KSL. [1,2,4,8,4,2,1] — that is the whole space, so there
is no projection error and what shows is exactly the accuracy of the local
exponential:
| old (EXPOKIT) | ttpy 2 | |
|---|---|---|
| 1e-3 | 1.33e-08 | 1.96e-15 |
| 1e-2 | 1.26e-05 | 2.09e-15 |
| 1e-1 | 6.71e-03 | 8.24e-15 |
The gap of 7–12 orders of magnitude has two causes: EXPOKIT is called there with
a fixed loose tolerance, and the order of the K- and S-steps in the backward
pass of the real branch of tt_ksl breaks the Strang palindrome (see above).
On the full manifold our version reproduces the dense exponential to machine
precision. Test:
tests/test_examples.py::test_ksl_is_exact_when_the_manifold_is_the_whole_space.
Separately: on an initial condition with unreachable ranks (
eigb.
| largest error | time | |
|---|---|---|
| old (PRIMME) | 9.0e-17 | 26.8 ms |
| ttpy 2 | 1.4e-16 | 14.0 ms |
So the replacement for PRIMME holds the same machine precision and is twice as
fast on this problem. Test:
tests/test_examples.py::test_eigb_matches_a_dense_symmetric_eigensolver.
- Fortran (
tt-fort,amen_f90,tt_f90,tt_eigb,tt_ksl,maxvol.f90,cross.f90),f2py,numpy.distutils, the submodules, thesetup.pybuild. Installation is an ordinarypy3-none-anywheel. - The binary tt-fort file format (
.tt). Old files have to be converted with the old package (or write a converter — the format is simple).
The signatures are unchanged: tt.eigb.eigb(A, y0, eps, rmax=150, nswp=20, max_full_size=1000, verb=1) -> (y, lam) and tt.ksl.ksl(A, y0, tau, verb=1, scheme='symm', space=8, rmax=2000, use_normest=1) -> y, tt.ksl.diag_ksl(...).
Old scripts run unchanged. What changed substantively:
-
The order of the K- and S-steps in the backward KSL pass is fixed. The
real Fortran branch (
tt_ksl) did K first and then S on every core, so the backward pass was not the exact reverse of the forward one and the Strang palindrome broke. The complex branch (ztt_ksl) had it right, and that is what is reproduced. The order of the scheme is verified numerically against an independent oracle:tests/test_verify_eigb_ksl.py::test_ksl_order_against_the_dense_projected_flowintegrates the projected ODE$y' = P_{T_y M} A y$ with a dense DOP853 (the projector is built from scratch inside the test, numpy only) and gives 1.00 forscheme='first'and 2.00 forscheme='symm', with a modelling error 25x larger than the splitting error being measured. The comparison has to be against the projected flow, not againstexpm(tau A) y0: relative to the latter the two schemes are indistinguishable. -
KSL measures and reports what a fixed rank cannot see.
check_rank=True(the default) computes the off-tangent part$(I - P_{T_y M}) A y$ and recordsdefect_relandstep_error_est(=$\tau |(I-P) A y| / |y|$ ) in the history; whenstep_error_est > defect_warnit raises an explicitRuntimeWarning. Measured:step_error_estpredicts the true step error againstscipy.linalg.expmto within 3%. It costs one TT matvec plus one sweep;check_rank=Falserestores the old price. -
The local exponential is our own Arnoldi with EXPOKIT-style adaptive
substepping instead of
dexp_mv;spaceis the same Krylov dimension, anduse_normestonly affects the choice of the first substep (and cannot change the result:test_ksl_knobs_do_not_change_the_answer). An unreachable accuracy is aRuntimeError, not a quiet answer. -
The local eigenproblem in
eigbatsize > max_full_sizeis solved byscipy.sparse.linalg.lobpcgon an implicit operator (instead of PRIMME); the true local residuals are measured and land inhistory.max_local_res. Problems smaller than$5B + 10$ always take the dense path. -
eigbmeasures its own residual and reports failure.ermax(the movement of the Ritz values) cannot tell convergence from being stuck: one-site ALS cannot grow a rank, so with$B = 1$ and a rank-1 initial guess the iteration stands still,ermaxdrops to 1e-14, and what used to come back waslam = 7.8e-3where the minimum is1.5e-4— silently. Nowcheck_residual=True(the default) computes$|A y_i - \lambda_i y_i|$ in the TT format (matvec + sum + QR sweep, never expanding into a dense vector) and puts it inhistory.res/history.res_rel; a relative residual aboveres_warn(1e-2) raises aRuntimeWarningwith the numbers. It costs about one sweep. The cure for the underlying problem is a higher-rank initial guess or$B > 1$ . -
sym_toldefaults to$\sqrt{\varepsilon}$ of the working dtype rather than a fixed1e-8: in float32 the projection of the local matrix is asymmetric at the 1e-7 level from rounding alone, and a fixed threshold rejected every float32 problem. -
New optional arguments (the default behaviour is unchanged):
return_history=Truereturns a history object (EigbHistory/KslHistory) with per-sweep or per-step records — written even atverb=0;eigbalso takeslobpcg_maxiter,sym_tol,check_residual,res_warn, andksltakeslocal_tol,check_rank,defect_warn. A non-symmetricAineigbis now an error rather than a silently symmetrized problem, andnswpwithout convergence raises aRuntimeWarningcarrying the indicator reached. -
The operator follows the vector's backend (the
amen_mvconvention): a numpyAwith a torch iterate used to die inside einops with aTypeError.
Four modules that were pure Python in the old ttpy as well. The signatures are
unchanged and the old import paths work (from tt.optimize import tt_min,
from tt.completion.als import ttSparseALS, from tt.riemannian import riemannian, from tt.solvers import GMRES, tt.min_tens, tt.min_func,
tt.GMRES); the implementation moved to tt/algs/{optimize,completion, riemannian,solvers}.py. All four gained an optional return_history=True
(ttSparseALS always returned its history) — records are kept even at verb=0.
- The left and right index sets are stored separately. In the old code one array
Jymeant the left set or the right one depending on the sweep direction; that worked, but it was never checked. - The sweep truncates the smoothed block to
rmaxsingular vectors in both directions. The old code went left through an SVD and right through a plain QR, so the index sets grew to about 4·rmaxon every other half-sweep. -
min_funccallsfunonly on a(P, d)array — including the final re-evaluation at the record point (the old code passed a(d,)vector there, and a vectorized function died at the very end of a successful run). - The returned value is always re-evaluated at the returned point;
history.consistencyis its discrepancy with what the sweep saw (nonzero only for a non-deterministic function). - New keyword arguments:
rho(the steepness of the default smoothing function$\pi/2 - \arctan((p - \lambda)/\rho)$ ;0.5formin_func, as in the signature, and1.0formin_tens, as in the old code),seed,return_history. -
history.evaluationscounts with repetitions: adjacent sweeps look at overlapping blocks.
- It no longer damages its input: the old code divided
cooP['values']by the norm in place. - The least-squares matrices are assembled for all samples at once by two
interface passes instead of one Python
getRowcall per (sample, slice) pair:$O(P d r^2)$ in BLAS instead of$O(P d^2 r^2)$ in the interpreter. - A slice that no sample touched keeps its previous value instead of being
zeroed (zeroing changes
Xwithout changing the functional — it silently destroys rank). -
alphais finally used (in the old code the call was commented out): it is thercondof the local problem, andalpha <= 0means the exact solution, the only mode with a guarantee that the functional decreases monotonically. -
convergednow means "the functional reachedtol". Stopping at an ALS stationary point abovetolisstop_reason='stalled'andconverged=False. Measured: a rank-2 tensor8x8x8x8(96 parameters), recovered at the true rank withalpha=0: with ~820 distinct samples 4 of 6 starts reachfit ~ 1e-15and 2 stall at1e-1; with ~1330 samples, 6 of 6. -
time.clock(removed in Python 3.8) is gone.
- The
numbabranch is gone: it duplicated the same mathematics with unrolled sixfold loops and only engaged when every rank in the listZwas equal. The same contractions througheinsumare faster and need no compiler. - They work on complex tensors as well (the frames enter the contractions conjugated); on real data the formulas coincide with the old ones.
- The
debug=Truebranch with its inlineasserts is gone — replaced by tests against a dense projector assembled from the SVDs of the unfoldings ofX, independently of this code. - They work on the torch backend (verified on CUDA).
- Restarts are a loop, not recursion (the old version called itself once per
restart and ran into the stack at a large
maxit). -
u_0is not damaged. The old version didu_0 += ...in place. - The true relative residual
$|b - A x| / |b|$ of the computedxis returned. The old version returned the residual of the first iteration of the last restart and declared convergence from it. - The small least-squares problem on the Hessenberg matrix is solved densely
(
lstsq) instead of with hand-accumulated Givens rotations: the old rotations were real and corrupted the complex case, and the inner product was conjugated on the wrong side. - The relaxation of the matvec accuracy is capped at one: a relative error of 1 means "return anything".
- Non-convergence is a
RuntimeWarningcarrying the residual reached, not silence. - The internal
_iterationargument (the recursion counter) is gone from the signature.
-
projectnow refuses to work at a rank-deficient point. Measured: a rank-1 tensor written with TT ranks(1, 2, 2, 1),$d = 3$ ,$n = 4$ , float64 — the formula returned a correct Hermitian idempotent projector (idempotence 1.6e-16) differing from the tangent projector at that point by 31 % of its norm. The previous guard ("orthogonalization changed the ranks") never fired: a QR never drops rank.X.round(0)does not reduce the rank either —chopateps <= 0returns the full size by definition;X.round(1e-14)does. The test is exact rather than heuristic: the singular values of the triangular factor$R_k$ of the left QR sweep, after the right orthogonalization, are exactly the singular values of the$(k+1)$ -st unfolding of$X$ . -
projector_splitting_addat such a point, by contrast, is left working: it measures 1.3e-15 on the same example, and refusing would be a regression.tt_qrthere gives orthogonality and reconstruction at 1e-15. -
ttSparseALSno longer stays quiet about complex data:cooP['values'](orx0) with an imaginary part is aTypeError, not a cast tofloat64behind aComplexWarningfollowed by "fit ~ 1e-30" for a fit to half the data. -
ttSparseALScounts and reports how many local systems the data fails to determine:info.underdetermined_slices,info.empty_slices,info.determined, plus aRuntimeWarning. Measured: a rank-4 tensor of shape6x6x6(144 parameters) from 38 samples givesfit = 4.9e-31,converged = True, and a relative error against the truth of 5.8. Nowdetermined = False.convergedwithoutdeterminedonly means "reproduces the samples". -
ttSparseALSscalesx0together with the data. Previouslymaxnsweeps = 0returned||values|| * x0, and "start from the exact solution" started from||values||times the solution. The factor goes into the zeroth core, which the very first local solve overwrites entirely — a run of one sweep or more is unaffected. -
min_tens/min_func:nswp < 1andrmax < 1are aValueError(previouslynswp=0died withAttributeError: 'NoneType' object has no attribute 'reshape');rmax=Noneworks as "no cap", as the docstring of_searchalways promised;history.index_sizesatd = 1is a list of pairs just as it is atd > 1(previously[1, 1], which mademax_index_setandrepr(history)die with aTypeError); an all-NaN block is aFloatingPointError, not anAttributeError. -
GMRES:eps < 0is aValueError(previously it silently meant "truncate nothing and never converge"); the division by an exact zero in the matvec accuracy relaxation, ateps = 0with an invariant Krylov subspace, is closed by an explicit branch. -
The oracle in
tests/test_ports.pywas wrong for the complex case, and it was the oracle that got fixed, not the code: the projector onto the row space of an unfolding was assembled fromvh[:r].conj().T, i.e. onto the complex conjugate of the row space. Such a matrix is Hermitian, idempotent and has the right trace, so no invariant sees it, and at a point of maximal rank — which is where the complex test stood (n = [3, 4, 3], rank 3, tangent space all 36 dimensions) — it is simply the identity. On a non-degenerate example (n = [3, 4, 5], rank 2, tangent space 24 inside 60) the discrepancy is 74 %. The correct oracle isvh[:r].T; it agrees with a basis of the tangent space built straight from the definition ($\mathrm{span}_k \tau(C_1, \ldots, dC_k, \ldots, C_d)$) to 2.6e-15, and withprojectto 6.7e-16. The tests that compare against the dense projector now check that the case is non-degenerate.
Not a legacy-compatibility matter (these arguments did not exist in ttpy 1.x),
but it changes the behaviour of code written against early 2.0 builds.
-
amen_solve(..., check_true_res=)is nowFalse. The exact residual$|A x - f| / |f|$ used to be computed after suitable sweeps and served as the stopping criterion. The product$A x$ has cores$(r^A_k r^x_k, n_k, r^A_{k+1} r^x_{k+1})$ — the ranks multiply — and on a preconditioned 2D problem ($r_A = 161$ ,$r_x = 122$ ) that is rank 19642 and 12.3 GB in a single core, with a measured peak of 91.8 GiB, for a number that is only printed. The same product rounded to 1e-12 has rank 256.Consequences for the caller:
info.true_resis nownan(not a guess but "not measured"), and the stopping criterion ismax_res, the local residual$|B_k x_k - \mathrm{rhs}_k|$ of every block before it is solved. The guarantee "never return an iterate worse than one already seen" now rests on an active measure. The exact residual is still available through the same argument and warns about its cost before allocating (true_res_budget). -
kslrefuses a stiff step instead of returning a number. The S-steps of the projector splitting run backwards in time, so for a dissipativeAthey amplify; the following K-step shrinks the data but not the rounding error that accumulated. Measured on$dy/dt = -(2^L+1)^2 \mathrm{Laplace} y$ : at$\tau|A| = 169$ it returned$|y|$ = 3.3e+106 where the exact norm is 0.307. Now every local exponential records its growth factor, the history carriesmax_growthandroundoff_floor, exhausting the digits entirely is aRuntimeError, and exceeding the requestedlocal_tolis a warning. The threshold is measured, not derived; the table of measurements sits next toKSL_GROWTH_EXPONENT. -
eigbwarns on the backward error, not on the relative residual. Theres_warnthreshold is no longer a fixed1e-2but$\sqrt{\varepsilon}$ , and it applies to$|A y - \lambda y| / |A|_2$ rather than$/ |A y|$ . The latter demands relative accuracy of every eigenvalue, which is unreachable at the bottom of the spectrum: onqlaplace_dd([10])the values are correct to 1e-9 absolute whileres/||Ay||is 3.0e-04.history.res_relis kept, andhistory.res_backandhistory.anormwere added. -
tt.permutereturns a compressed representation. Bubble transpositions used to leave rank slack behind: on a three-peak separable function shuffled into Morton order at$d = 15$ , rank 1024 for a tensor whose own rank is 102. The tensor was right, the representation was not.