Skip to content

OptimusPy 2.0.0 release candidate - #84

Merged
nicolasbisurgi merged 87 commits into
cubewise-code:optimuspy2dot0from
nicolasbisurgi:optimuspy2dot0
Sep 29, 2026
Merged

nicolasbisurgi merged 87 commits into
cubewise-code:optimuspy2dot0from
nicolasbisurgi:optimuspy2dot0

Conversation

@nicolasbisurgi

Copy link
Copy Markdown
Collaborator

This brings optimuspy2dot0 up to the 2.0.0 release candidate. After this merge, what's left is the merge to master and the v2.0.0 tag.

What's in it

  • The optimizer. It searches and selects inside the cube's storage dimension frame.
    • dimension_position_rules lock a dimension to the position they name, and are applied before the search starts.
    • VMM/VMT is always put back, including when the cube config is rejected.
    • A run with no RAM signal is reported as such, rather than as a result.
  • Optimize DB. A plan-first pass that reorders every cube on an instance by leaf-element count, fewest first, within a time limit.
    • A run can resume, and it re-enables chores on every exit.
    • Already optimized cubes (a storage order that differs from the presentation order) are skipped unless include_optimized is set.
    • Every run writes an HTML report next to its plan and run files.
  • The web UI.
    • Navigation: a Home page with the getting-started guide; Reports organised by instance, cube and run, with Build report for older Optimize DB runs; and sidebar links to the documentation, the website and GitHub.
    • Jobs: one job runner over a replayable event log, so a job survives a reload or a second tab. Stop never aborts a rebuild in progress, and the Activity Monitor follows every running job.
    • Settings: config.ini is linked from where it lives (for example RushTI's) or copied in, and never edited; the CLI uses the same file. The folders for saved cube configs and Sync Order exports can be chosen.
    • Sync Order: confirms before applying, skips cubes whose order already matches, and reports each cube.
    • Hardening: the server answers only its own page and never sends stored secrets back, and TM1 errors reach the page as TM1's message.
  • The CLI and the bundles.
    • The executable opens the web UI with ui or a double-click.
    • An unknown instance is a one-line error.
    • config/settings.ini holds the linked config.ini, the UI port and whether a browser opens.
    • CI builds the bundles for Windows, Linux x86-64 and Linux arm64, runs each executable before keeping it, and ships it with the example config and samples.
  • Documentation.
    • A User Guide, with a mode chooser and step-by-step pages, and docs rewritten to match what the code does.
    • Real screenshots of the 2.0.0 UI and report.
    • A light landing page, the README, and the release notes in CHANGELOG.md.

Checks

  • python -m pytest -q: 453 passed. The 10 live tests are deselected by default.
  • python -m mkdocs build --strict: clean.
  • CI on the fork: Tests and Build Executable (Windows, Linux x86-64, Linux arm64) pass. The Windows bundle and the Linux arm64 bundle were each run on a real machine.

After the merge

  • The public site redeploys from this branch. cubewise-code.github.io/optimus-py is already built from optimuspy2dot0, so this merge publishes the current docs and landing page there.
  • No release yet: Build Executable keeps the bundles as run artifacts, and only master publishes a release.
  • The merge to master: tag it v2.0.0 (a 2.0 tag from 2021 already exists), date the ## 2.0.0 heading in the changelog, and remove optimuspy2dot0 from the triggers in build.yml, deploy-site.yml and test.yml.

🤖 Generated with Claude Code

nicolasbisurgi and others added 30 commits September 18, 2026 22:09
Adds `optimize_db.py` — the heuristic pass that sweeps an instance cube by
cube against a wall-clock budget, writing a resumable run artifact as it
goes — together with the wiring that already referenced it:

- `cli.py`: the `optimize-db` subcommand and the `tm1_connector` factory
- `core.py`: `_dimension_metadata` extraction + cache, which
  `optimize_db.py` imports (`_collect_dimension_metadata`'s signature
  changed with it)
- `ui.py`: `start_optimize_db_job` and its four handlers
- `static/`: the Optimize DB page
- `mkdocs.yml`, `README.md`, `CONTEXT.md`, `docs/`: nav entry and prose

Until now the module was untracked while three tracked files already
imported or navigated to it, so a fresh clone could not import
`optimuspy.cli` or build the docs, and 56 of the 154 tests were absent.

Also ignores PLAN.md, pending.txt and logs/ — local scratch that nothing
in .gitignore covered.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two defects that between them broke the resumable-overnight-sweep promise.

`execute_plan` initialised `status = "completed"`. Every legitimate exit
sets `status` explicitly, but the `except Exception` guard re-raises an
unexpected error — a KeyError, a TM1py payload change — without touching
it, so the `finally` persisted the stale "completed" into the run
artifact. `optimize_db(resume_plan_id=...)` then reported "already
completed — nothing to resume" and a sweep that died at cube 40 of 300
was recorded as a success and lost. The initialiser is now "failed";
only the `for...else` can say completed.

`prepare_resume` wiped `run["samples"]`. With the throughput window gone
`fits_in_budget` returns `(True, None)` for the first cube after any
resume, so a run resumed ten minutes before its inherited deadline could
start a six-hour rebuild that has no safe abort. The samples now survive
the resume; prior throughput on the same instance is still valid.

Both are covered end to end through the scripted FakeServer, which grows
a `probe_raises_once` hook for an error the module never anticipated. The
budget test sits with the other resume tests in the executor file rather
than the planner file, which has no TM1 fake.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
.github/workflows/ held only build.yml (PyInstaller, master only) and
deploy-site.yml — nothing ran pytest anywhere. The 56 Optimize DB tests
that just landed protected nothing without a gate.

Runs `python -m pytest -q` on master and optimuspy2dot0 for both ends of
requires-python, 3.9 and 3.12, so the declared floor is actually tested.
The install is non-editable so the job also proves the package builds and
its dependencies resolve.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every executor was constructed with the cube's presentation order
(get_dimension_names) while the server enforces, and the run restores,
the storage order (get_storage_dimension_order). The engine permuted one
list and was judged against another; on a cube whose two orders differ,
a position index meant one slot to OptimusPy and another to TM1.

Pass the storage order to all five executors, collect dimension metadata
over it, and resolve optimize_position / validate optimize_dimension
against it. Nothing in _execute_optimize_mode needed the presentation
order afterwards, so the read is gone; the UI and scan paths keep their
own get_dimension_names calls, which are genuinely presentational.

Rename measure_dimension_only_numeric to last_slot_locked, inverted at
the point of computation so callers read positively: it was always the
lock condition (`not measure_dimension_only_numeric`) wearing the name of
its own negation. Log one INFO when the lock is detected.

Behaviour change: on a cube whose presentation and storage orders differ,
an existing "optimize_position": 3 now targets the third *storage* slot.
That is the intended meaning and is documented in task 8.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ge frame

The comment claimed the string_dims set is authoritative "NOT
presentation-order [-1], which need not be the string dim". Since the
executors moved into the storage frame, self.dimensions[-1] IS the locked
dimension, so the claim inverts the truth and reads as if the string_dims
indirection were still load-bearing.

Describe what the block actually does now: the lookup and the relocation
it feeds are a redundant second signal from the old presentation frame,
and for a cube with one string dim the relocation branch cannot fire.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Seven places enforced the string-element constraint, each with its own
idea of what the constraint is and which order it applies to. This adds
the one module that will replace them: built from the cube's storage
order plus whether the last slot is locked, it answers whether a
candidate order is admissible and returns the reason when it is not.

The lock is a server constraint, so it binds every order source. The
user preferences it also carries — position rules, excluded dimensions,
ignored orders — are supplied by the greedy folds alone; a caller that
supplies none gets the lock and nothing else.

Pure: no TM1Service, no I/O, no logging. The caller logs the reason. That
keeps it offline-testable in the same category as tau.py, which is what
lets the scripted fold tests stop needing a stand-in server.

_unsatisfied_position_rule is _violates_position_rules moved verbatim,
defect included, so that routing the folds through the frame is a pure
refactor. Its correction lands in its own commit.

Nothing calls the frame yet.

Adds **order frame** and **locked slot** to CONTEXT.md under a new "Order
admissibility" heading.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The predicate reported a rule as violated when the dimension WAS at its
configured position, and the caller skips on a violation — so the feature
forbade exactly the layout the user asked to lock. Its integer branch
also compared int(pos) - 1, making positions 1-based, while
docs/advanced/dimension-position-rules.md documents 0-based.

A rule is now satisfied when the dimension is at the position it names,
and integer positions are read 0-based as documented. 'first' and 'last'
are still accepted for the end slots. Nobody can be relying on the old
behaviour: it was the inverse of both the documentation and the intent,
and it shipped with no tests.

required_index resolves a rule's position to a real slot or to None, and
None means the rule names no slot and is passed over. That covers an
absent dimension, a word that is not a keyword, an index the cube does
not have (too large or negative), a float — truncating 2.7 to slot 2
would silently reinterpret a malformed config — and a bool, which would
otherwise resolve to slot 1 because bool subclasses int.

Passing these over matters as much as inverting the predicate: once
inverted, a rule that names no slot would make EVERY candidate order
unsatisfiable, so the greedy would evaluate nothing and the run would
report success. Ignoring them is also the pre-existing behaviour. Turning
them into the errors the docs promise is startup validation's job, where
the message can name the config field.

The refusal message now names the required position as well as the actual
one, since it is what a DEBUG-level skip line has to explain.

Behaviour is unchanged for users until the folds are routed through the
frame — MainExecutor still carries its own copy of the predicate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The "order frame" entry described a two-tier model — server constraint
versus user preference — written before the frame gained its
well-formedness check. A malformed order is neither of those: it is a
config error, and it is the one tier that fails loudly rather than
skipping.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fold A froze the string dims out of the pool and relocated them to the
back; Fold B's seed sorted them last and its refine list excluded them;
_string_last_skip guarded the last slot a third time; and
_greedy_skip_permutation carried its own copy of the position-rule
predicate. All four are gone. One call to the frame now covers the
locked slot, the ignored orders and the position rules, and
movable_dimensions/reserved_positions give both folds one answer about
what may move and which slots are closed.

This is a behaviour change, not dead-code removal. The relocation block
appended EVERY dimension in string_dims to the end, not just the locked
one — its own comment conceded "TM1 permits at most one such dim, but a
list is handled defensively". Dimensions are shared between cubes, so a
cube can legitimately contain a dimension carrying string elements from
another cube's use while not being this cube's measure. The old code
moved it to the back; the locked-slot rule leaves it placed by
cardinality like anything else. That is the one case where the old and
new rules genuinely disagree, and both folds now have a test for it.

dimensions_to_exclude, orders_to_ignore and dimension_position_rules
leave MainExecutor with string_dims: the frame owns them, and a second
copy on the executor is a standing invitation to enforce them an eighth
way. cube_dim_number goes too — it existed only for _string_last_skip.

A refusal is logged at DEBUG with its reason and counted by code, and
the counts are reported once per cube after the run (decision 8). A
malformed candidate raises instead of being skipped: the greedy builds
every order by swapping within the cube's own dimensions, so one could
only mean a bug in the sweep.

_has_string_elements survives for the two targeted executors until they
get the frame.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Predefined orders had no enforcement at all; set mode called
update_storage_dimension_order directly; the position and dimension
optimizers each asked the server, per candidate, whether a dimension had
string elements. All four now consult the frame, and the three tiers are
handled as they differ rather than as if they were one thing.

Tier 1, well-formedness. Predefined orders are checked in full BEFORE
the first reorder, so a typo cannot fail a sweep half-way with cubes
already modified, and every bad entry is reported rather than the first.
validate_cube_config settles what needs no server — shape, empties,
duplicates — so those never cost a TM1 connection to discover. Set mode
logs the offending name and exits non-zero: a TI process must not see
success after a typo silently did nothing.

Tier 2, the locked slot. A predefined order that moves the locked
dimension is skipped with its reason and the run continues. Set mode
warns, skips the reorder and exits 0 (decision 6), with one greppable
line saying the cube is unchanged and that the exit code means nothing
was applied — the log is the only channel ExecuteCommand leaves.

_has_string_elements is deleted. It ran get_element_types inside a
sweep, once per candidate, to re-ask a question the frame answers once
per cube. It was also wrong: it asked whether the CANDIDATE carried
strings and let a numeric one into the last slot, displacing the locked
dimension, which TM1 would then have rejected.

Consequences worth stating: on a locked cube, optimize_position "last"
and optimize_dimension naming the locked dimension now evaluate nothing
at all. That is the honest answer — the slot has one legal occupant —
and the skip counts say so in the run summary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A malformed entry in predefined_orders, or one naming a dimension the cube
does not have, is a config error. Both checks raise ValueError —
validate_cube_config before the connection, _validate_predefined_orders from
inside the run — and neither sat inside an except, so the operator got ten
lines of Python stack trace with the message buried in it.

Tier 1 exists because the log line is the only channel a TI process calling
this via ExecuteCommand has; a stack trace is not that line. The set/optimize
branch now catches (ValueError, FileNotFoundError) exactly as the optimize-db
branch at cli.py:186-188 already does, printing "ERROR: <message>" and
returning 1 — the style CONTEXT.md documents as "exit 1, no traceback".

The exit code does not change: it was 1 before and is 1 now. Only the channel
does. A malformed cube config JSON now reports the same way as a bonus, since
json.JSONDecodeError is a ValueError.

The catch is as wide as optimize-db's, so it also swallows a ValueError raised
from deeper in the run. The traceback stays reachable for those: it is logged
at DEBUG with exc_info.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fold A computed mid = int(len(dimension_pool)/2) — a pool that excludes the
user's dimensions_to_exclude and the locked dimension — while Fold B computed
it from the full-length resulting_order. mid decides which half of the cube a
position is in, and ADR-0002 keys pruning strength to the metric that owns a
position: the back half is RAM-ranked and pruned at full TAU_RAM, the front
half is query-ranked and pruned at the wider TAU_QUERY, or left unpruned on a
process cube.

Deriving it from the pool made a position's *metric* depend on how much of the
cube the run happened to be searching. Excluding one dimension from a 6-dim
cube moved the split from 3 to 2, so slot 3 — a front, query-ranked position on
the same cube without the exclusion — became a back position pruned at
TAU_RAM. That is the outcome ADR-0002 explicitly rejects.

Both folds now call tau.midpoint(len(self.dimensions)). Fold B's value is
unchanged; only Fold A's moves.

Positions Fold A sweeps, before -> after (the locked slot is skipped by the
occupant guard in both, so it is absent from the swept list):

  6 dims, locked                [0, 4, 1, 3] -> [0, 4, 1]
  6 dims, 1 excluded            [5, 0, 4, 1, 3] -> [5, 0, 4, 1]
  8 dims, locked                [0, 6, 1, 5, 2, 4] -> [0, 6, 1, 5, 2]
  7 dims, locked, 2 excluded    [0, 5, 1, 4] -> [0, 5, 1, 4, 2]
  7 dims, locked                unchanged: [0, 5, 1, 4, 2]
  5 dims, locked                unchanged: [0, 3, 1]

The direction is not uniform: where the pool midpoint sat below the cube's own
midpoint the sweep shortens, and where two or more exclusions pushed it further
down it lengthens. Either way the swept set now depends only on the number of
dimensions, so two runs over the same cube with different exclusions no longer
search different halves of it. Anyone comparing run durations across this
release should expect a locked even-dimension cube to get one position faster.

Note this does NOT undo the extra position Task 3 introduced on a 7-dimension
locked cube ([5, 0, 4, 1] -> [0, 5, 1, 4, 2]). Seven is odd, so the pool
midpoint and the full midpoint were already equal there; that sweep was widened
by the removal of the freeze-shortened range, not by mid, and the wider sweep
is the correct half-split.

tau.ranking_for_position and tau.tau_for_position are untouched — the bug was
entirely in what they were fed.

Tests assert the ranking of every position, not only the swept ones: Fold A
skips a position whose occupant has drifted out of the movable pool, and that
skipping can mask a moved split. All four discrimination tests fail against the
old derivation.

Refs: docs/adr/0002-cardinality-pruning-keyed-to-optimization-metric.md

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
configure_logging hardcoded level=logging.INFO and none of the CLI's eleven
arguments touched the log level, so every DEBUG line in the codebase was not
suppressed-by-default but unreachable at any flag.

That was already true before this branch; this branch made it matter. The order
frame reports every refusal at DEBUG — the locked slot, an ignored order, a
position rule — so "why was my order skipped?" had no answer available. Two
shipped docs pages promised output nobody could turn on:

  docs/advanced/dimension-position-rules.md — "suppressed by default"
  docs/concepts/string-element-constraint.md — linked to an "Optimization
  Logging" section of how-it-works.md that did not exist

And 6fade52 justified its wide except (ValueError, FileNotFoundError) with a
DEBUG traceback that could not be read, which made the mitigation fictional.
With -v it is real, so that catch stays as written; there is now a test that
fails if the traceback stops being recoverable.

configure_logging takes verbose and sets the root level explicitly after
basicConfig, which is a no-op once the root logger has handlers — otherwise -v
would depend on what configured logging first. The stdout handler is added only
if one is not already attached. The call moved below parse_args, which is the
only way the flag can reach it; nothing above that line logs.

Docs: how-it-works.md gains the "Optimization logging" section the other two
pages link to, giving the INFO skip count, the -v skip lines, and the three
greppable reason codes. Both links now point at it by anchor.
string-element-constraint.md's example message was also the pre-frame one; it
is updated to what the code emits, since a doc telling operators to grep for a
string that no longer exists is worse than no example. The rest of that page is
Task 8's.

Tests deliberately avoid caplog: that fixture forces the root level, so a test
written around it would pass whether or not -v worked — the same class of
non-existent mitigation this commit removes. A handler that respects the root
level is used instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tion

Two behaviours the documentation has always described did not exist: startup
validation, and pre-application. Both land here, in the frame.

Pre-application belongs in the frame, not beside it. A rule pinning a dimension
to a slot makes that slot un-sweepable and that dimension un-movable — the same
two facts movable_dimensions() and reserved_positions() already answered for
the locked dim and the excluded dims. They now answer it for a third reason as
well, and OrderFrame.pre_applied_order() seats the pinned dimensions while the
rest keep their relative storage sequence. There is no second list in
MainExecutor: enforcing a preference in a second place is how the seven copies
this branch deletes came about.

Both folds now start from the pre-applied order. Fold B's seed builds on it,
and Fold A's resulting_order starts from it instead of self.dimensions.

That last one carries a consequence worth stating. When pre-application moves
something, the measured original order is no longer the order Fold A is
standing on, so its "keep what is here" option at every position would carry a
RAM figure for a different order. Fold A now measures its starting point when
it differs from the storage order, as Fold B has always measured its seed. That
is one extra reorder, and only when the rules actually move something.

_unsatisfied_position_rule reads the resolved pins rather than re-reading the
raw rules, so a rule becomes a slot in exactly one place. With the slots
reserved the greedy no longer generates a violating order at all; the check
stays as the backstop that says so.

Validation turns every case 5a deferred into the error the docs promise —
absent dimension, unreadable position, out-of-range index, float, quoted "3" —
plus two rules on one position, two rules on one dimension, and a collision
with the locked slot in either direction. It reports every problem at once and
runs before any reorder is sent.

A type slip is not reported as a typo. "3" and 3.0 name slot 3 unambiguously
and are refused only because the documented type is an integer, so the message
says write 3, not "3". An out-of-range value reports the range even when it
arrived as a string, so "99" does not cost two round trips. "middle", None,
True and 2.7 get the generic message, which is what they deserve.

One judgment call beyond the brief: a dimension named in both
dimension_position_rules and dimensions_to_exclude is now an error. One says
leave it where it is and the other says move it to a specific slot, and
pre-application would have silently resolved that in favour of the rule.

Not changed, and worth knowing: resolve_position (core.py:198), which reads the
separate optimize_position field, is still 1-based. dimension_position_rules is
0-based per its own documentation. The two fields disagree; that is outside
this task.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The split between offline and live tests existed only informally: tests/ made
zero TM1 connections while samples/validate_v11_v12_parity.py and
samples/smoke_resume_drop.py were live tests that pytest never collected.

Register the marker and default to -m "not live", so `pytest` on a fresh clone
or a GitHub-hosted runner is offline by construction rather than by the accident
of nobody having written a live test yet. Live runs are opt-in:

    pytest -m live --instance tm1srv01     # v11
    pytest -m live --instance tm1srv02     # v12

conftest gains the plumbing those runs need: --instance / --tm1-config, and
session-scoped tm1 / is_v12 fixtures built from the real config.ini. There is
deliberately no fake behind them — an unreachable server skips the test rather
than passing it against a simulation.

test_suite_split.py asserts the default selection offline, because a live test
that leaks into the default run does not fail fast; it opens a socket on a
runner with no route to the host and times out.

.github/workflows/test.yml is unchanged and stays offline-only: it runs bare
`pytest -q`, which now carries the exclusion from pyproject.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Neither changes what is asserted.

test_fold_b.py seeded fold B's resume state with a "seed_order" key. No code
writes it and _run_fold_b reads only current_order and pass_index, so the key
described a checkpoint schema that has never existed — the kind of detail that
gets copied into the next resume test as if it were required.

test_seed_numeric_measure_regression.py called its fixture PRESENTATION_ORDER.
Executors are handed the storage order since Task 1, so the name now points at
the list the engine does not use. Renamed to STORAGE_ORDER; the assertion it
guards — the seed is cardinality-ascending and never pins a numeric measure last
— survives the frame change intact and is untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…etire the scripted evaluator

install_scripted_evaluator replaced _evaluate_permutation whole, so every
offline test ran against a hand-built PermutationResult. That meant the helper
owned a second copy of the %-chain — it re-derived the percentage from
context.current_ram, which is the very thing PermutationResult computes — and it
had grown its own regression test. A test helper with a regression test is a
fake.

Split the one TM1 call out of _evaluate_permutation into _measure_permutation,
returning a Measurement of what the server reported. Everything else stays where
it was: the pending write, the reanchor one-shot, the %-chain, the run artifact,
the progress log. An offline test now replaces only the measurement and gets
real PermutationResults back, so all of that is exercised rather than simulated
— install_offline_measurements supplies the numbers and nothing else.

The percentage it reports is the change against the order the cube was in a
moment ago, which is what update_storage_dimension_order returns; the first
reading is absolute because there is nothing for a percentage to be relative to.
test_the_percent_chain_derives_ram_from_the_servers_percentage now tests
production's chain instead of the helper's inverse of it.

Four hand-rolled executor factories (object.__new__ plus a list of attributes)
were all missing _reanchor_needed, invisible for as long as the tests also
replaced the method that reads it. They are replaced by one offline_executor()
that calls the production constructor with tm1=None, so a sweep that reaches for
the server fails loudly instead of passing quietly.

test_fold_a_freezes_excluded_dim set ex.dimensions_to_exclude on the executor,
which nothing reads — exclusions are the frame's. It passed because "Excl" had
no cardinality entry and the tau frontier pruned it anyway. It now goes through
the frame with a cardinality inside the undecided cluster, and fails if the
exclusion is removed.

test_resume_integration.py keeps its FakeTM1 and gains a note saying why: crash
recovery reads the cube back from outside the executor, so there is no seam to
supply numbers through. It is the suite's only fake and is marked as not a
precedent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
metrics.py is the RAM truth source for both TM1 versions and had 170 lines and
zero tests. Its row-reading half is pure and now runs in CI (test_metrics.py):
the conversion table, and that an unknown Unit fails loud rather than passing a
raw number through — the failure mode there is silent, because a number that is
1024x wrong still looks like a cube size.

The half that only a server can answer is live (test_live_metrics.py): which
Unit each version actually reports (B on v11, KB on v12), that a settled cube
reads the same size twice — the premise the v12 plateau loop rests on — and that
ram_source_ready leaves the v11 Performance Monitor as it found it.

samples/smoke_resume_drop.py and samples/validate_v11_v12_parity.py were live
tests that pytest never collected, because testpaths is ["tests"]. They are now
driven by test_live_resume_drop.py and test_live_parity.py. The scenarios stay
in samples/, called rather than copied: they are also the thing you run by hand
against an instance that is misbehaving, and a second copy would drift.

The crash-and-resume pair in test_checkpoint_folds.py stays offline. It touches
no server — the "crash" is a raised exception and the checkpoint manager is a
local recorder — and what it proves (a resumed fold skips the locked positions
and still recommends the uninterrupted run's best order) is arithmetic that
deserves to run in CI. The server half of that scenario, where the cube is
actually left when a connection dies mid-reorder, is what the promoted smoke
test now covers live, on both branches.

    pytest                                  287 offline, 10 live deselected
    pytest -m live --instance tm1srv01      v11
    pytest -m live --instance tm1srv02      v12
    pytest -m live --v11 tm1srv01 --v12 tm1srv02    the cross-version gate

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The "do not stage" list from docs/superpowers/plans/2026-09-18-land-optimize-db.md
had no mechanical backstop, so a `git add .` swept all of it into a test-refactor
commit. Naming the paths here makes the convention enforceable rather than
remembered.

parity_snapshot_v2.json joins the run-artifact block next to parity.json; the
rest are local scratch. The patterns are anchored because `images/` at the repo
root is tracked and only its icon.png is a stray duplicate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…he tree

"the suite's ONE fake" was false and written as a standing instruction, so the
first maintainer to grep for Fake would have found it contradicted on the first
hit and discounted the rest. The tree also holds six in test_optimize_db_executor.py
and a two-method reorder stub in each of test_inflight_status.py and
test_ram_reanchor.py.

The note now names them, says why each is deferred rather than blessed, and keeps
the one rule that is true: nothing new joins the list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The docs described the rule OptimusPy used to implement: place every
string-bearing dimension last, relocating if needed. It now locks the last
storage slot when the dimension there has strings, and never relocates.

concepts/string-element-constraint.md is rewritten around that (nav label and H1
become "The Locked Slot"; the filename stays so inbound links hold). It states
the rule once, spells out that it keys off the POSITION rather than the
dimension — the shared-dimension case where v1 and v2 genuinely differ — and
gives the per-mode behaviour and the real log lines.

modes/position-optimization.md: positions resolve against the storage order, and
targeting the locked slot now evaluates nothing rather than sending TM1 a reorder
it rejects. modes/predefined-orders.md and modes/set-mode.md: the tier-1/tier-2
split, including why a skipped reorder exits 0 and a malformed order does not.

Every quoted message is copied from a real run, not paraphrased: the first draft
of the predefined tier-1 example invented an error text the code does not emit.

Also corrects two pages that still asserted the old rule in passing:
concepts/cardinality-aware-greedy.md said any string-bearing dimension is locked
last, and modes/optimize-db.md's link text. optimize-db's own behaviour is
unchanged — it still applies the old rule, deliberately deferred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ADR-0002 and ADR-0003 both assume the arrangement this work replaced. A fourth
ADR is cheaper than amending two, and rewriting a decided record to match later
code destroys the reason the record exists.

The decision: the engine reasons about get_storage_dimension_order() and nothing
else, one pure module answers admissibility for every order source, and the last
slot is locked only when the dimension in it carries strings — never relocated.

Three findings are recorded because they are invisible in the diff and will be
misread as bugs:

- The shared-dimension disagreement. The old rule moved every string-bearing
  dimension to the back; the locked-slot rule places a non-last one by
  cardinality. This is the one case where old and new genuinely differ, and it is
  why the deletion was safe. core.py:745 and optimize_db.py:202 still apply the
  old rule, so the greedy and the suggestion now diverge on such a cube —
  deferred deliberately, and stated here so it is not met as a contradiction.
- The _has_string_elements defect: it asked whether the CANDIDATE carried
  strings, so a numeric one entered the locked slot and TM1 rejected the write.
  Consequence recorded too — optimize_position "last" on a locked cube now
  evaluates nothing, which is the honest answer for a one-occupant slot.
- The mid unification. Claim worded precisely: the range and the break point
  depend only on the dimension count, so the tau ranking per position no longer
  shifts under dimensions_to_exclude. NOT that the swept set is exclusion-
  independent — it is not. ADR-0002 is referenced, not amended. Before/after
  swept-slot table included; the direction is not uniform.

The three admissibility tiers are not restated — they are CONTEXT.md's
vocabulary (01f3eac) and the ADR points there.

Line citations into deleted code were checked against the commits that removed
it rather than copied from the plan, which had them stale by three commits.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…pass

Both live fixtures could leave the OptimusPy_Parity_Test cube and its eight
dimensions on a shared instance.

test_live_parity.py built BOTH instances outside its try, so a v12 setup failing
after v11 had already built its cube left v11's behind. Setup moves into a module
fixture's try/finally, each instance is torn down independently so one failure
does not skip the other, and a teardown that itself fails now fails the test —
a leftover cube on a shared server should not be a warning under a green run.

test_live_resume_drop.py used a yield fixture, which covers a failing test but
not a setup_instance that dies half-built: cube created, dimensions partly
created, yield never reached. Wrapped in try/finally. teardown_instance guards
every delete with exists(), so it is safe on a partial or absent build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The tm1 fixture skipped on any exception, its own comment naming "bad
credentials" among them. Running the live suite against tm1srv02 produced
"10 skipped, 287 deselected" — visually identical to a laptop with no TM1
access, and in fact an HTTP 401 from a server that answered in 45ms.

A run that tested nothing must not look like a run that was never asked to.
These ten tests are the justification for coverage that left CI, so a stale
password gets an error naming the instance and the status code. An unreachable
host still skips: that is the machine having no route, not a claim about the
code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The stabilization note claimed a mitigation the observed failure mode walks
straight through. The loop detects a gauge that is RISING; it cannot detect one
FLAT AT THE WRONG VALUE, because two equal reads are the same observation
whether the cube settled at its true size or the gauge is stuck on the pre-load
skeleton. Its window is also shorter than the ~5 minute lag seen after a
300k-cell load.

Removes "The winner is derived from %-deltas and is correct regardless" with it.
That sentence made a frozen gauge look harmless, and a live run contradicts it:
if the percentages were real a wrong baseline would scale every figure by one
constant and leave both the ranking and ram_reduction (a ratio) intact. Fifteen
identical RAM figures means every update_storage_dimension_order returned 0% —
the %-channel, OptimusPy's only measurement channel, was dead too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
When every candidate comes back 0.00% the %-chain never moves, every row in the
report carries the same RAM, and a winner is still picked — by whatever broke
the tie — with nothing in the output saying the search had no RAM measurement to
go on. Observed on v12 against a cube whose cube_memory_used gauge was still
reporting the pre-load skeleton: fifteen permutations, one distinct value.

ram_signal_is_dead is a post-run check over data the run already holds: no extra
server call, no new read. core logs it at WARNING before the winner is announced
so the two are read together, plus the known v12 cause when applicable.

Tests both conditions rather than just the percentages, so a resume re-anchor —
0% everywhere but two distinct RAM values — is not reported as this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a tie

Four fixes to the cross-version gate, from a live run that failed on greedy for
reasons that were entirely the harness's.

Settle gate. setup_instance now reads the gauge on the empty cube it just
created, then blocks after the fill until the reading is clear of that skeleton,
and raises if it never is. The threshold is measured rather than guessed. A gate
that cannot measure its own fixture has nothing to say about parity.

Ties. The settle gate removes the artifact but not the tie underneath it: Dim6
and Dim7 measured 0 bytes apart on both servers, and at a cardinality ratio of
2.08 they are inside TAU_RAM, so the greedy is meant to settle them by
measurement. Demanding a total order where the data supports an equivalence
class would flake again on a green run. compare now asks each version about the
OTHER version's winner using its own numbers, and calls it a tie only if both
say there is no measurable difference. An order a version never evaluated is an
unknown, not an equal. This asserts the thing that is actually true.

No adoption. _build_dimensions/_build_cube used update_or_create, so a fixture
left by a crashed run was silently adopted. Worse than it looks: a crash after a
mode run leaves the cube REORDERED, and the next run reads that back as its
"original order" — the value the gate compares across versions. setup_instance
drops a leftover instead. The resume-drop fixture still reuses a warm cube, but
now checks it is the fixture before trusting it.

Transport vs disagreement. core.main swallows its fatal errors and returns
False, so the report could not tell a dropped socket from a version difference —
a ConnectionResetError(54) got diagnosed twice. Mode runs now capture the logged
ERROR, and the report labels the outcome ERROR (not FAIL), names the side that
died, quotes the reason and says a reset is a harness failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…seline

The stabilization note described cube_memory_used as a sampled gauge that lags
about five minutes after a load on v12. That was wrong in every part. Measured
side by side on 11.8.02200.2 and 12.6.4 against the same cube and the same six
rearrangements, v12 tracks a reorder to the byte within about a second; it is
v11 whose gauge sits one sample behind its own reported percentages. The repo
already held the disproof in samples/parity_snapshot_v2.json — a passing
six-mode parity run whose v12 greedy found a 37% reduction it could only have
found by measuring, with a 40,960 B cold baseline sample sitting right beside
it.

What actually goes wrong is not version-specific and is one cause, not two: a
cube whose data is not resident costs the same in every order, so the skeletal
baseline AND the 0% from every reorder are both truthful answers to a question
worth nothing. They are correlated, which is the part the old note got
backwards by implying the %-chain could route around a bad baseline.

So the plateau loop keeps its job (it is sound against a rising gauge) and
stops claiming a mitigation it never had, and the case it cannot see is caught
by arithmetic instead: cube_num_populated_*_cells arrive in the same by_cube()
payload and report correctly while cube_memory_used is still skeletal.
September measured 40,960 B against 300,000 populated cells — 0.137 bytes per
cell, which no storage engine can produce, against ~224 B/cell for a real
reading on that cube. Readings below 1 byte per populated cell are now refused
rather than optimised against.

The floor is far below any plausible density on purpose: this rejects readings
that are impossible, not readings that are surprising. Absent or zero cell
counts get no verdict, so an empty cube and a server that does not report the
counts are both left alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ram_signal_is_dead was right to exist and its message was pointed at the wrong
culprit. It named v12's pre-load skeleton as "the usual cause", which reads as
a claim about the server version — and the repo holds a passing six-mode
parity run on v12 that disproves it.

The finding it reports is narrower and stands: every order came back 0.00%, so
the run had nothing to rank and its winner is a tie-break. That is a statement
about THIS CUBE on THIS RUN. The cause is residency, not a version: an
unmaterialised cube costs the same in every order, so 0.00% everywhere is the
server answering honestly.

The docstring also now records what stopped an earlier reading of this signal
from being believed — a 0% for one rearrangement is ordinary. Not every
rearrangement costs memory; an identity, an adjacent swap of the two smallest
dimensions and a swap of the two largest all return 0% on v11 and v12 alike.
Only ALL of them coming back zero means the run learned nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The post-fill gate waited for cube_memory_used to climb clear of the empty-cube
skeleton. That tested a symptom, and it tested the wrong one: a cold sample is
normal and appears in a passing parity run. It also could not pass on a server
whose gauge is correct but whose cube is not yet resident, which is the case
that actually broke September's run.

Replaced with a probe of the channel every mode run depends on. Apply one
rearrangement, require a non-zero percentage back, restore. The rearrangement
is chosen because it is KNOWN to cost memory on both engines rather than
assumed to: moving the largest dimension to position 6 measured -12.5004% on
11.8.02200.2 and -12.4931% on 12.6.4 in the same controlled run.

That choice is the whole value of the probe, so it is pinned by tests. Most
rearrangements of this fixture are free on BOTH engines — the identity, the
D6<->D7 swap and the D1<->D2 swap all return 0% on each — and a probe built on
one of those would pass against a cube that is not resident and prove nothing.
It is also what made an earlier investigation conclude v12 had no RAM signal:
four orders, picked for being the smallest perturbations available, are not a
sample.

The cube is restored on every path, including a raising one. A probe that left
the fixture rearranged would have corrupted the original order the gate
compares across versions — the same class of bug the no-adopt rule fixed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
nicolasbisurgi and others added 29 commits September 24, 2026 18:07
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…the log is always in logs/

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d selection

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d-artifact v7

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…he diagrams

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…no --include-optimized

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… and the web UI

The hero animates an illustrative greedy run on a sales cube: dimensions
reorder, memory drops, each view runs five times and the median counts,
and orders that aren't better are put back. TM1 v12 (powered by TM1py's
MetricService), the web UI and Optimize DB each get their own section,
and the rest of 2.0 sits in tiles you can browse. The palette is light
and built on the logo's blue and amber, and the logo is cropped to its
artwork so it renders at its real size.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…site speaks in one voice

Facts corrected against the code: exit code 2 is a usage error, the winner
is picked by a tolerance band on every metric (process time included), the
locked slot is a position rather than a dimension, the greedy prints no
"of N", results land under results/<instance>/, the set-mode sample has no
views, Sync Order jobs live in the tab that started them, and the v2 page
no longer shows a set config the code rejects.

Structure: site_url carries /docs/ like the deploy does, the banner pages
have a real H1, the Requirements pages become Design Notes at the end of
the nav, the Optimize DB page gets its own UI doc, sections follow the
sidebar and the mode hierarchy, and What's New is built on the changelog.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… CLI

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lot can be picked

Add Rule offers Lock to Position with slots numbered from 1, like the Overview
table, and writes them from 0 as JSON numbers. One rule per dimension and one
per slot, as the run requires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lan and the Folders card

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The same instance name can point at a different server in the new file,
so its cached cube list and dimension data no longer apply.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gs page

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ized is set

A cube whose storage order differs from its presentation order was
ordered on purpose, often by a measured optimization, and the
leaf-count pass would overwrite it. The plan skips it as
already_optimized; include_optimized (the Include optimized box on the
Optimize DB page) reorders it like any other cube.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ebsite

Home keeps its instance tiles and the Tips & Getting Started guide after
you connect, and opens with the guide expanded. The sidebar footer links
the documentation, the landing page and GitHub; Home links the first two.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilicon

The Linux job runs on ubuntu-latest and ubuntu-24.04-arm; the arm64 build
is optimuspy-linux-arm64.tar.gz. The install page says which to pick.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Optimize DB run

Every Optimize DB run ends with optdb_report_<plan-id>.html beside its plan
and run files, written from those two files alone on every way the run ends,
and again when a resumed run ends. It shares the single-cube report's page
shell and stylesheet: the stat cards, the note that the saving shows after a
TM1 restart, a chart of memory saved per cube, the per-cube table with its
orders, the skip ledger, the chores and the run settings. A report that can't
be written is a warning and never changes the run's outcome. The CLI prints
the report's path after the run summary.

The Reports page lists every file under results/ by instance, then cube, then
run, with Optimize DB runs under their instance, plans never run as Plan only,
and anything else under Other files. Each run opens its report, with its data
files as small links. POST /api/optimize-db/report builds a missing report
without TM1, for Build report and for Report under Previous runs. JSON data
files are served as JSON. #/results opens Reports, and the Optimize page's
per-cube tab shows the cube's runs in the same rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…aving reads in MB

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…onitor clears when a job ends

View Prior Reports is enabled once the cube has runs on the connected
instance. While a job runs, the Activity Monitor asks the server again
every few seconds, so it clears when the job ends even if no page is
following that job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@nicolasbisurgi
nicolasbisurgi merged commit 2340a8c into cubewise-code:optimuspy2dot0 Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant