O4d: sync junior evidence, sampler, workflow, and calibration fixes - #208
Merged
Merged
Conversation
…ound
Adversarial review of the first two commits. The correctness pass called them not
mergeable; the test pass got 19 mutations past the suite. Both were right, and the
second one located the pattern: everything asserted by reading source text rather
than running it was weak.
The export decision is now RIFT/misc/distance_grid.distance_grid_inputs, a function
that returns (distance, ln_weights, notes, warning). It was six statements inside
analyze_event, which nothing can import, so the only available tests read the source
for names. Nine separate ways of getting it wrong left those tests green: inverting
`if _resampled`, which turns the whole retained-set change off; reading the reserve
and then re-reading _rvs anyway; exporting psi as distance; zeroing the weights
before the check; never printing the warning. The driver is now the call plus the
printing, and the decision is tested by calling it.
Fixed, each with a test that fails without it:
* force_peak never reached AV or the portfolio, the two samplers that actually
emit a .dgrid -- they build their reserve directly, so they got the default.
Rather than a second sampling rule, the exporter now divides the forced row by
its own inclusion probability, and n_subsample is recorded so it can. Measured
on 8000 rows capped at 800 with a dominant row: raw ESS 3.6, corrected 87,
population 99. Dropping the row instead reports 506, because the subsample
usually misses it.
* --skymap-file registers ("declination","right_ascension") as one parameter, so
_rvs[key] is (2,N), the vstack went ragged, and those runs silently got no
reserve at all. Tuple keys are flattened to their component coordinates.
* The weight was re-derived term by term, which RvsRecord.log_weights() documents
as wrong: a row with a zero prior and a zero sampling prior gives NaN, and one
NaN takes the whole export to .dgrid.skipped. One conjunctive keep-mask now.
* The exporter reached past _rvs_record_for to sampler._warm_seed_reserve exactly
when the record declines to describe the rows. Measured: exported peak at
1914 Mpc for rows peaking near 700. It falls back to _rvs instead.
* The resolution check went silent on the collapsed pass it exists for: the
integrand column saturates, the weights come back uniform, ESS 4000 of 4000
while the curve was flat to 0.13 nats across a true 1130-nat span. It now also
reads the sampler's own n_eff, which no weight column can contradict.
Also: per-site AST sweeps replace text searches that a comment satisfied and that
let mcsamplerGPU's second fair-draw site pass on the first site's call; the
per-backend integrand_is_log flags had no coverage and now must match their own
RvsRecord; two existing CI tests that sliced fixed source windows are parsed
instead; the moved _rvs read carries a verdict in the fair-draw ledger; and the
core-unit floors go to 487/475, measured.
One regression this suite could not see: the refactor deleted params_out and the
distance-prior columns, every test stayed green, and a real ILE run failed at the
very end with UnboundLocalError. Nothing executes the driver, so the export block
now gets an order-sensitive name-resolution check.
ILE-GPU-Paper demo, ldas-pcdev13 GPU 3, --fairdraw-extrinsic-output-n-max 50:
GMM 50 rows / 50 distinct, lnZ 58.72011 = .dat lnL, warned at sampler n_eff 5.9;
AV 50 rows / 50 distinct, lnZ 58.95728 = .dat lnL.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Base moved under the PR (noise-evidence work, #176). Only .travis/test-core-units.sh conflicted: both sides raised the core-unit floors. Resolved by MEASURING the merged tree rather than adding the two deltas -- 494 collected / 482 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…-safe-toy-lnL tests: make the L0-rescue toy integrands backend-agnostic
…e repo root test_mcsampler_foridiots.py, test_mcsamplerEnsemble.py and test_mcsamplerEnsemble_AdaptationDemo.py contain no test functions. They are demo scripts whose work happens at module level, so pytest ran them at collection time and they dropped test_fig_0.png, cdf.png, my_cdf_sampler.png and cdf_adapt_test.png into whatever directory pytest was invoked from -- usually the repo root, where a `git add -A` sweeps them into the commit. Renamed them to demo_*.py, matching the convention already used by demo_mcsampler_NF.py and demo_mcsampler_portfolio_with_oracle.py in the same directory. pytest's default python_files patterns (test_*.py, *_test.py) no longer match, so it does not import them; running them by hand is unchanged. The figures are the point of these scripts, so they are kept as-is; tmp_path was not an option because it is a fixture and there are no test functions to take it. Nothing referenced the three by name outside Code/test/README.md, updated here. Also ignore figures at the repo root as a backstop for running a demo from there. The rules are anchored with a leading "/", so they cannot reach the 19 tracked .png files, all of which live under MonteCarloMarginalizeCode/ or docs/source/images/. Verified on ldas-grid with the CVMFS igwn python: the same directory collection that writes three PNGs on a050a24 writes none here, and pytest ignores the renamed files when they sit beside a real test. Tracked .png count still 19. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…s cached
save_intg turns on only when tempering_exp > 0, but the adaptation block runs for
any n_adapt > 0. At the default tempering_exp = 0 it read a cache that was never
filled: integrate() raised NameError('int_vals'), a name that has never existed in
that scope, and integrate_log() raised KeyError('log_weights'). Both are reachable
from add_parameter(adaptive_sampling=True) plus integrate(). The README's
"easy-to-read" demo died on its first mcsamplerGPU call.
Each branch now uses the current chunk's weights when no history exists.
ci_roster.txt guessed fval**tempering_exp for the missing expression; int_val
(fval*prior/p_s) is used instead, matching the sibling branch two lines down. Over
5 seeds on the demo's Gaussian at nmax=5e4, int_val gives n_eff 941+-300 and answer
scatter 5.3e-4, against 69+-3 and 3.3e-3 for raw fval; all arms unbiased to <1%.
Points are sliced to the depth the weights actually reach rather than to n_history.
The two differ in the chunk-only branch, and for a sampler reused for a second
integrate(), where _rvs[p] carries over while "integrand" restarts. That pair
reached numpy.bincount as a length mismatch.
Renamed the demo to demo_*.py: with the crash fixed pytest would import it, run it
for 30 s at collection, and drop test_fig_0.png in the repo root. Corrected its
mcsamplerGPU reference value, which compared a normalized-prior integral against
sqrt(2 pi) sigma and so printed a factor-2.5 "PROBLEM".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ci-roster-check failed on the rename: .travis/ci_roster.txt still listed the three files under their old test_* paths, and a roster entry for a path that no longer exists is a silent no-op, so the census demands the lines go. The census walks test_*.py and *_test.py only -- the same patterns pytest uses -- so demo_*.py is outside its scope, as the pre-existing demo_mcsampler_NF.py and demo_mcsampler_portfolio_with_oracle.py already are. The three entries are therefore dropped rather than repathed. All three were HANDRUN, never gated, so no coverage changes. The foridiots entry carried a note worth keeping: the demo dies on an undefined `int_vals` in RIFT/integrators/mcsamplerGPU.py, in the `not save_intg` branch of integrate() -- core sampler code, not the demo's, and reachable on CPU. That note no longer belongs in a roster the file is not in, so it moves into the demo's own header, where the next person to run it will meet it. Still unfixed here, for the reason the original note gave: guessing the intended expression is RO's call, not a side effect of test hygiene. The line number is deliberately omitted -- the roster's 1324 had already drifted to 1360. Short headers added to all three demos saying they are not pytest targets and which figures they write. test-ci-roster.py now passes: 260 test files, 212 gated, 48 rostered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…view's findings
Third adversarial review, of the remediation delta. Verdict was mergeable with two named
fixes; both are here, plus the cheap ones it listed.
THE PRIOR NOW COMES BACK WITH THE ROWS. distance_grid_inputs returns ln_prior_at_rows,
evaluated inside it. The caller used to evaluate it two lines later, which left a gap in
which the prior could be taken at the fair draw's distances while the curve was built from
the retained set's. Measured: that moves the exported lnL by 2.87 nats and turns
ln_prior_d_sampling from a d^2 ramp into a flat column, with the whole suite green. Same
test-invisible shape as the params_out deletion, so the fix is structural rather than one
more assertion.
The tuple test asserted shapes, not values: reshape(-1, 2).T has the same shape and
interleaves declination with right ascension, and the warm-seed rescue builds a covariance
from those columns. It now checks each component against what was handed in.
mcsamplerPortfolio builds its reserve through the shared adapter instead of its own vstack,
which is where the ragged-tuple and NaN-weight defects still lived. Two implementations of
one thing drift; a test pins the one.
Smaller, each measured:
* _ess over n EXACTLY equal weights returned n*(1-eps), so a flat weight set read as
"fewer than one effective sample per bin" at n = 3, 5, 6, 8, 9, 10 and not at 2, 4, 7,
11. Rounded to the representable neighbour.
* A NaN sampler n_eff passed the new clause silently, because NaN < n is False. The
predicate is now `not (n_eff >= n_rows)`: a pass that could not measure its own
convergence is the case the clause exists for.
* A capped, peak-forced reserve with no recorded n_subsample cannot have its forced row
divided by an inclusion probability, so it is declined rather than used uncorrected.
* The remedy line told users to raise --fairdraw-extrinsic-output-n-max on the
retained-set path, where it does nothing. It is now path-aware, and reports the
achieved n_eff again.
Two gates were lying, and widening them is how I found out:
* audit_backend_contracts probed the literal "make_warm_seed_reserve(", so it recorded
builds_reserve=False for mcsampler, mcsamplerGPU, mcsamplerEnsemble and mcsamplerNFlow
and stayed green while THIS BRANCH gave all four a reserve. The probe now asks for any
of the three builders, and the four flips are re-recorded.
* The core-unit floors must be the plugin-free counts. I first set them from the junit
line, which includes 3 pytest-subtests, pinning the gate to an environment that
happens to carry that plugin. Now 497/485, with the ladder entry the file asks for.
Correcting commit 3: --skymap-file's tuple registration is dead in this driver
(integrate_likelihood_extrinsic_batchmode:2180 is `if False:`), so no .dgrid run was
losing its reserve to it. The flattening is still right, and the live tuple registrations
are in integrate_likelihood_extrinsic and the LISA driver.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… the verdict
mcsamplerAdaptiveVolume.update_sampling_prior_selfish evaluated the integrand on the device
array with no try/except and no fallback. That method is how mcsamplerPortfolio drives a
VARAHA member, so a host-only integrand -- a CI toy, a benchmark, a user's CPU likelihood --
died on any cupy-importable host with "Implicit conversion to a NumPy array is not allowed",
and passed wherever cupy is absent, which is every CI runner.
It now evaluates device-first and retries on a host copy, remembering the verdict in
_integrand_wants_host, as mcsamplerAdaptiveVolume.integrate_log, mcsamplerPortfolio and
mcsamplerNFlow already did. The host copy is the MODULE-level identity_convert, the
counterpart of the identity_convert_togpu that put rv on the device three lines above; both
read module-level cupy_ok rather than self.xpy.
This does NOT make the contract universal, and an earlier draft of this message wrongly said
it did. mcsamplerGPU.py:774 still evaluates with no fallback and no flag, though it is not in
the same trap because it draws with self.xpy instead of force-pushing with the module-level
converter, so it never hands a device array to a host-configured run. mcsamplerEnsemble.py:200
does the opposite, converting to host before every call, so a device-native likelihood under
--sampler-method GMM always receives numpy. Both are out of scope here; neither is a
regression from this change.
mcsamplerPortfolio hands the member the verdict it already has, since its own evaluation runs
earlier in the same chunk. That is speed, not correctness: the member retries on its own. It
only sets the flag, because a portfolio that has not learned "host" is indistinguishable from
one whose target is device-native.
Two guards the propagation pays for. A propagated verdict can be WRONG -- the portfolio may
have latched on a transient failure of a device-native likelihood -- so the branch that reads
it unlatches and returns to the device instead of dying on an unguarded host call. And both
AV sites now short-circuit on `not cupy_ok`: without it, where cupy is absent identity_convert
is the identity, so an integrand that legitimately raises was invoked a second time with the
byte-identical array before the exception propagated. Measured at 2 invocations on ldas-grid;
doubled side effects, and doubled RNG draws for a likelihood that marginalizes distance or
calibration internally, which moves a seeded run.
Measured on the file list of the ci.yml job this PR edits, plus
test/integrators/test_portfolio_{gmm_member_trains,restrict_and_warm}.py, base 0be696f vs
this branch:
ldas-grid, no cupy 319 passed, 0 failed -> 330 passed, 0 failed
ldas-pcdev2, RTX 3080, cupy 12.0.0 310 passed, 14 failed -> 323 passed, 11 failed
The three repaired on the GPU host are test_fairdraw_on_the_cupy_backend and the two tests of
test_portfolio_gmm_member_trains.py. The eleven that remain fail identically at the base
commit and are untouched here: six in test_portfolio_fairdraw_backend.py, two in
test_gmm_truncated_score.py, three in test_portfolio_restrict_and_warm.py. Two root causes,
both filed separately: the fairdraw suite's _to_host stand-in is monkeypatched over the real
cupy.asnumpy and then declines to convert real cupy arrays, and gaussian_mixture_model
._normalize allocates on the device from a module global while writing host samples into it.
test_av_selfish_host_integrand.py runs the device path on a host without cupy, using a
stand-in that numpy dispatches to through __array_function__/__array_ufunc__ and that refuses
implicit conversion with cupy's message; where cupy is real the same tests run against it.
Ten mutations revert-checked in an isolated worktree with an unmutated control through the
same scoring path: all ten caught on ldas-grid, nine on ldas-pcdev2, where dropping the
no-cupy short circuit is invisible by construction because the defect it guards exists only
where cupy is absent and its test is skipped there.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…st-fallback AV: give update_sampling_prior_selfish the host fallback every other call site has
…m-limits-o4d ILE: derive the psi and phi_orb priors from their sampling ranges (evidence change: -ln 2)
…no-root-pngs test: stop pytest running three demo scripts that write figures to the repo root
…e cupy is real
The _DeviceArray stand-in gave these tests reach on a CPU-only host, but two
harness assumptions only held where cupy was ABSENT, so five tests failed on
every cupy-importable host and passed everywhere else.
* `_to_host` converted only _DeviceArray. It is monkeypatched over
PF.identity_convert, which on a GPU host IS cupy.asnumpy -- so the patch
disabled the real conversion rather than extending it, and the genuinely
device-backed member draws (mcsamplerEnsemble's self.xpy is cupy whatever
the portfolio's is) collided with the host aggregation arrays in
mcsamplerPortfolio.draw. It now does cupy.asnumpy's job too.
* `np.random.seed` does not seed the run. The Ensemble member draws from
cupy's RNG. Measured on ldas-pcdev2 (RTX 3080, cupy 12.0.0): three repeats
of the SAME arm, same seed, gave lnZ -5.960285685330888 /
-5.904156714565715 / -5.954741248760214, a 0.056 nat spread -- wider than
the between-arm difference test_the_fix_does_not_move_the_integral was
flagging, so that bit-identical assertion was reading noise. Seeding cupy
as well, four repeats agreed to the last bit, so the assertion stands
unchanged rather than being given a tolerance.
Also corrects the docstring's host-portability claim, and pins the _to_host
contract directly so a future narrowing names its own cause.
Production code is untouched. Verified on both backends at f61286f:
ldas-grid (no cupy) 7 passed 2 skipped, ldas-pcdev2 (CUDA_VISIBLE_DEVICES=0)
9 passed, against 6 failed 2 passed there before. Reverting both halves of the
fair-draw fix under this harness fails 7/7 on ldas-grid and 8/8 on ldas-pcdev2,
so the tests still detect what they were written for.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Base moved again: the AV selfish-host fallback (#342), the psi/phi_orb prior derivation (#306), and the two chips this work spawned -- backend-agnostic L0 toy integrands (#340) and the demo scripts that were writing PNGs to the repo root (#343). Auto-merged with no conflicts. Verified rather than assumed: both sides survive in all five overlapping files, git diff and git diff -w agree line for line so there is no reindentation damage, every file parses, and the count-pinned core-unit floors still hold (500 collected / 488 passed, plugin-free 497/485, which is what they are set to). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rdraw-gpu-harness portfolio fair-draw tests: make the host-conversion harness work where cupy is real
…ate-bins .dgrid: build the distance grid from the retained set, not the fair-draw export resample
…PU-adapt-no-intg-cache # Conflicts: # MonteCarloMarginalizeCode/Code/test/README.md
test_mcsampler_rosenbrock.py calls optparse parse_args() at module level. pytest imports it during collection, it sees pytest's own argv, and exits. That raises out of the collection loop as an INTERNALERROR, which --continue-on-collection-errors does not catch, so collection of the directory stops there. Every file sorting after it was never collected: 16 items before, 99 after. The file has no test functions. Like the demos already named demo_mcsampler_NF.py and demo_mcsampler_portfolio_with_oracle.py beside it, its work happens at import and it writes fairdraw_rosenbrock_*.dat and cdf_rosenbrock.png into the current directory. Renamed to demo_mcsampler_rosenbrock.py, which pytest's python_files patterns no longer match. Running it by hand is unchanged. Dropped its ci_roster.txt line rather than repointing it: the census walks only test_*.py, so an entry naming the renamed file reports as a deleted-file no-op. Re-adding the line fails test-ci-roster.py with exit 1. Also updated its usage comment and the entry in Code/test/README.md; nothing else referenced it by name, and test_integrator_studies.py does not list it. Verified on ldas-grid with the CVMFS igwn python. The command in the report now exits 1 with no INTERNALERROR and 99 collected. The 5 remaining collection errors are the same 5 on the unmodified base, all rostered LEGACY or OPTDEP. test-ci-roster.py passes: 263 test files to 262, rostered 51 to 50, HANDRUN 10 to 9. test/test_mcsamplerEnsemble_extended.py has the same defect and still aborts a collection of the parent test/ directory. It is left alone here because .travis/test-integrate.sh runs it by name as a script, so a rename has to move with that gate. No CI job collects a directory, so neither case is CI-red today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The exported rows are already an equal-weight fair draw, so the weight both resamplers form, exp(lnL - lnLmax) * (p/ps) / Npts, has to be constant over them. alpha2/alpha3 = 1.0 tilted them by the likelihood a second time (measured spread ratio 0.7020 against 1/sqrt(2)). Write p = 1, ps = exp(lnL - logZ), and set sample_n so Npts is this file's row count rather than a process-wide counter's running total. A missing P/fiducial_epoch/logZ/sigma_lnL now skips the XML with a stderr warning instead of raising, which had taken test_jax_fairdraw_export.py from 34 passed to 22 failed. Adds the 5-D --phase-marginalization test that a surviving mutant exposed, and ILE's 24 short option forms. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
integrate_log's fallback used log_weights, which at tempering_exp = 0 is
0*lnL + ln p - ln p_s and carries no likelihood at all, so the histogram replayed the
proposal's own noise. It now uses log_integrand, the log-space twin of the int_val
integrate() uses. Over 8 seeds, on a proposal tilted 10:1 away from a target at +0.6,
log_integrand concentrates the adapted proposal by at least 178x on every seed;
log_weights by as little as 0.02x.
The reuse path traded a loud failure for a silent one. The end-of-pass cleanup
reindexes every _rvs key with one index list built from the integrand length, which on
a carried-over parameter record picks its first rows, pairing each likelihood with
another draw's parameters. Before this branch that state died at bincount instead.
_trim_rvs_to_record drops the rows the record does not describe, and is a no-op
whenever the lengths already agree.
A degenerate chunk now leaves the proposal alone instead of writing nan into the
histogram, which surfaced far away as an out-of-bounds index in pdf_from_hist. That is
GPU-only, because the fval.sum()==0 skip above is gated if not(cupy_ok). The sum stays
on the device; only the scalar used for the check comes back to the host.
mcsampler.py carried the same defect, masked because n_adapt defaults to 0 there and
to 1000*n in mcsamplerGPU. With n_adapt=2 at the default tempering_exp it raised
KeyError('integrand') and now integrates unbiased. The other four integrators are clean.
The tests asserted only that nothing crashed: weighting by fval, by a flat vector, or
by log_weights all passed them. They now measure the weight rule. One adaptation step
on a constant integrand must return a 10:1 proposal to the uniform prior, which
importance weights do (tilt 0.97 to 1.02) and the raw integrand does not (3.38 to 3.51).
Six mutants fail now, including the three that used to pass.
Also renamed test_mcsampler_gpu.py, a second demo of the same shape. Its roster entry
said LEGACY "imports ourio", but ourio imports fine: the int_vals crash was what made
the file unimportable, so fixing that turned a stale reason into a CI failure.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
adapt_weight_exponent is the argparse dest in the ILE drivers, which rename it to tempering_exp before calling a sampler. No integrator reads the argparse spelling, so the EOS driver's weight exponent was swallowed by **kwargs and every sampler used its own default. The intrinsic driver had the same bug until f9f4456; this driver was forked from it and missed the fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
samplers.adaptive_volume_sample passes igrand_fairdraw_samples=False for AV, so the integrator keeps the retained weighted population and its per-row log_joint_prior / log_joint_s_prior. Return both, carry them through analyze_one, and publish them as alpha2/alpha3 the way the classic ILE does: exp(lnL)*p/ps downstream is then the true importance weight with nothing cancelling. On that path the XML is the weighted cloud while the .dat stays a fair draw, because util_ConvertJAXILEFairdraws.py reads the sidecar as equal weight. Under portfolio, nuts and map the integrator consumed the weights and log_weight is None, so those rows keep the p = 1, ps = exp(lnL - logZ) pair that makes the resampling step a no-op. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…brock-demo test: stop one demo script aborting collection of test/integrators
…adapt-no-intg-cache mcsamplerGPU: adapt on the current chunk when no integrand history was cached
test_jax_ile_extrinsic_xml_export.py grew from 18 tests to 35. Measured on ldas-grid (import cupy fails there, numpy backend), IGWN conda python 3.11 / lal 7.7.1: per-file 455 over 43 files, junit 458 collected / 446 passed / 12 skipped / 0 failed. The junit counts carry 3 pytest-subtests entries, so the floors stay pinned to the plugin-free 455/443. The gate PASSED at the old 438/426, which is the under-coverage this roster exists to catch rather than a reason to leave it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…extrinsic_xml # Conflicts: # .travis/test-core-units.sh
test_mcsamplerEnsemble_extended.py and test_vector_coordinates.py are named test_* but are scripts. Each runs an optparse parse_args() at module level. pytest imports one during collection, it parses pytest's own argv and exits, and that raises out of the collection loop as an INTERNALERROR. The whole run stops there, so every file sorting after it is never collected. --continue-on-collection-errors does not catch it. On Code/test/: 2278 items collected before, 3015 after. An AST sweep of every test_*.py under Code/ found eight files that parse argv at module level. Only these two reach parse_args; the other six fail an import first, which is an ordinary collection error and already rostered LEGACY. #344 renamed the demos in integrators/ out of the test_ namespace. That fix does not apply here. Both files are live gates invoked by name with --as-test, from .travis/test-coord.sh and .travis/test-integrate.sh, and test-ci-roster.py keys reachability off the file name. Calling a gate demo_* would also misdescribe it. So they keep their names and conftest.py tells pytest to skip them. An ignore list can go stale in silence, so conftest.py re-checks its own entries and fails collection with a UsageError if an ignored file grows a test function or stops existing. Both branches were driven on a scratch copy, where the unmutated copy still collects 6 tests, so the check can pass as well as fail. Emptying collect_ignore reproduces the base: rc 3, 2278 collected, 124 INTERNALERROR lines. Verified on ldas-grid with the CVMFS igwn python. Collection of Code/test/ goes from rc 3 with an INTERNALERROR to rc 1 with none. The 31 files that now report collection errors were each confirmed to error the same way on the unmodified base, so they are unmasked rather than caused. test/integrators/ is unchanged at 110. Both gate scripts still exit 0 when run the way their .travis scripts run them, and test-ci-roster.py passes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
RIFT's own Gaussian mixture (RIFT/integrators/gaussian_mixture_model.py) set
self.xpy = xpy_default in __init__, with no way to override it. So on any host where cupy
imports, the class became device-backed regardless of what the caller was using.
_normalize then allocated `out = self.xpy.empty((n,d))` on the device and wrote host
samples into it, which cupy refuses:
ValueError: non-scalar numpy.ndarray cannot be used for fill (core.pyx:737)
Passing plain numpy is the ordinary way to call this class. So fit(), update() and score()
rejected ordinary use on every GPU host, and passed where cupy is absent, which is every
CI runner. Same family as junior PR #342.
The fix is inference, not PR #342's device-first-then-retry. A retry is the right shape
when the unknown is someone else's callable: you cannot ask a likelihood which backend it
wants, you can only try. Here the backend is written on the argument. fit(X) is told by X,
and guessing wrong then catching the exception is a slower way to read the same fact.
Inference also reaches the half a retry cannot. score() handed a host caller a DEVICE
array back. That is not an exception anywhere, just a wrong type downstream.
One rule, three sources, in this order. A method given an array reads THAT array. That
covers _normalize, _unnormalize, score, the estimator's EM, _xpy_logsumexp,
_mixture_log_density_normalized and fit_gmm_adaptive. fit() and update() bind self.xpy
from the samples they fitted, so the model remembers. sample(), the defensive component,
pruning and the merge have no argument at all, so they read the model's own parameters
through _model_backend. That is how a hand-built host model, never fitted, still draws
host samples. xpy_default survives only as the starting guess for an unfitted model.
gmm and estimator now take an xpy= override.
Three further module-global reads went with it, all the same shape. _xpy_logsumexp keyed
on cupy_ok, so it pushed a host argument onto the device and returned a device result. The
gpu_logpdf-vs-scipy dispatch in _e_step, score and _mixture_log_density_normalized also
keyed on cupy_ok. That would hand scipy a device array, or send host arrays through the
device routine. It now keys on the resolved backend. No path that works today changes
which routine it calls, or its numbers. identity_convert and identity_convert_togpu stay
as instance attributes, but bind to self.xpy rather than to the globals, so a host-fitted
model cannot push its own parameters onto the device. Internal call sites use _to_host and
_to_backend, which need no such binding.
MEASURED, base 2f9d7f7 vs this commit, on the four suites that reach this module
(test_portfolio_restrict_and_warm 22, test_gmm_adaptive 6, test_portfolio_gmm_member_trains
2, test_gmm_truncated_score 2 = 32 tests), plus the 31 added here:
ldas-grid, no cupy 32 passed, 0 failed -> 63 passed, 0 failed
ldas-pcdev2, RTX 3080 cc86, cupy 12.0.0 27 passed, 5 failed -> 63 passed, 0 failed
ldas-pcdev13, RTX 2080 Ti cc75, cupy 12.0.0 27 passed, 5 failed -> 63 passed, 0 failed
The same five fail on both GPU hosts:
the three in test_portfolio_restrict_and_warm.py that PR #342's message filed, plus both
tests of test_gmm_truncated_score.py. All five raise the fill ValueError out of _normalize.
gmm.print_params() was broken on the same hosts and no test covers it. On a device-fitted
model it raised TypeError("Unsupported type <class 'numpy.ndarray'>") at base, and returns
without error here. Across the wider integrator/portfolio list on ldas-pcdev2, 6 failed /
399 passed at base becomes 1 failed / 435 passed. The survivor is
test_seeding_reproducibility.py::test_deterministic_gpu_histogram_is_bit_reproducible,
which fails identically at base and is untouched.
The DEVICE path is bit-identical to base. md5 of means, covariances, weights, score(),
sample(), the parameters after update(), and fit_gmm_adaptive all match, as does the
container type of `means`. Nothing that worked before computes a different number.
An adversarial review before un-drafting found a REGRESSION this change had introduced.
_merge coerced self.weights with a bare np.asarray. self.weights is not always the float
array fit() leaves behind. mcsamplerEnsemble.create_wide_single_component_prior assigns
the python list [1], and that is live code, reached from
integrate_likelihood_extrinsic_batchmode. np.asarray of a list of ints has dtype int64, so
`self.weights[i] = weight` truncated every merged weight to 0 and took score() to its
1e-300 floor for every sample, with no error. Measured: weights [0, 0], score [1e-300]*5.
The same review found _merge rebuilding `means` unconditionally, which turned the (k,d)
array fit() leaves behind into a list of (d,) arrays on every update. Same numbers,
different type, no diff line to notice it by. Both are fixed and both now have a test.
The second is why the container type is pinned above.
test_gmm_backend_dispatch.py runs the device path on a host without cupy. It stands in for
a device array with a class numpy dispatches to through __array_function__ and
__array_ufunc__; a raising __array__ alone is not enough. That stand-in raises cupy's fill
ValueError when a host array is written into a slice of it, and refuses to compute
alongside a host array the way cupy does. An array module allocates it. Where cupy is real
the same tests run against it.
Beyond the defect the suite carries three things. A control that the fix did not force
everything onto the host: a device-fitted model must keep returning device arrays, because
MonteCarloEnsemble._sample writes model.sample(n) straight into a self.xpy array. A
call-site spy on gpu_logpdf, because the two density routines return the same numbers on
every case that works, so only which one ran distinguishes them. And an assertion on the
component count fit_gmm_adaptive chooses. Every candidate fit there is wrapped in
`except Exception: continue`, with a fall back to k_min. A backend mistake inside that
ladder does not raise. It returns a worse proposal.
Thirty-four mutations revert-checked in an isolated worktree, with an unmutated control
through the same scoring path, control clean before and after: 33 caught on ldas-grid.
The survivor drops gmm.fit()'s coercion of log_sample_weights. estimator.fit() coerces
them again from its own samples, so that mutation preserves behaviour. The line stays
because it holds a local invariant that does not depend on what estimator.fit does, but no
test can distinguish it.
The battery reached 34 through four rounds, each of which found the suite weaker than the
round before. An 18-mutation set scored 17. Widening it to 34 dropped that to 29, and the
gap was always the same mechanism: the stand-in permitted mixed-backend operations that
cupy refuses, so a backend mistake still produced a plausible number on a cupy-free runner.
It now refuses them in all three spellings -- the operator, the array protocols, and the
module namespace -- and returns 0-d device arrays from indexing and reductions rather than
letting them decay to host scalars. Two mutations, an unconverted `bounds` in _normalize
and an unconverted `weights` in score(), were caught only by that last part. The suite also
gained an absolute anchor in d=2 with unequal box widths, because every other test compares
the two backends to each other, and a numerical mutation moves both identically: prod->sum
on the box volume survived everything until that anchor existed.
The suite seeds both global generators so it is deterministic, and restores them so the
next suite in the same pytest process is unaffected. An earlier revision seeded and walked
away. That is the same leak with a different value, and it broke
test_seeding_public_paths.py::test_skymap_oracle_draws_do_not_depend_on_string_hash_salt
downstream, on a GPU host only, because the cupy branch is the only one that runs there.
cupy.random.seed mutates the live RandomState in place, so the fixture swaps in a fresh one
and puts the original object back. Measured on ldas-pcdev2: both streams are byte-identical
whether or not the suite has run.
Most of the module-global reads were observable only on a model built by ASSIGNMENT rather
than by fit(). A fitted model agrees with itself, since self.xpy and the parameters' own
backend say the same thing. Hence the hand-built-model tests.
calmarg/extrinsic_handoff.py's reconstruct_gmm converts every stored array onto the
device. Its docstring said `bounds` was the reason. That has changed. `bounds` is now
converted per call, and `means` sets the backend. Dropping the conversion makes sample()
return host arrays, MonteCarloEnsemble._sample raise, and MonteCarloEnsemble.integrate
catch it and _reset() every gmm_dict entry. The run then continues COLD, with the seed
discarded rather than reported. The docstring now says that, and names which conversion
must not be missed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
test-core-units.sh 543/531 and test-jax.sh 898, measured after merging rift_O4d rather than carried forward. Neither side's raise described the merged tree: 543 is 29 ABOVE 497+17, because the base added test files beyond the two named in its own roster entry, and 898 is 2 above 868+28 for the same reason. Summing floors would have pinned both gates below their real coverage while still passing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ble-extended test: stop two CI gate scripts aborting collection of Code/test
Adversarial review of the rename found that my_exp is now live, so a negative
value is no longer inert. The ratio divides by max(Y), which carries lnL_shift,
while the historical guard tests the unshifted Y_orig: --lnL-shift-prevent-overflow
above the peak lnL yields a negative exponent, which adapts AWAY from the peak
(measured n_eff ~1 and lnZ wrong by 7.6 nats on a toy where the fix gives ~1000).
Test holes, each now covered by a mutation:
- the spy binds its own extra_args, so a colliding key in the driver's real
extra_args passed 8/8 (either spelling: one is silently dead, one is a
run-time TypeError). Checked by AST instead, and the docstring's claim to
cover it is corrected.
- the per-module substring scan passed with one entry point's read hardcoded.
Now checked per integrate/integrate_log, with unconditional delegation to a
sibling allowed; mcsamplerGPU.integrate delegates only under use_lnL, so
walking the whole body would have exempted it.
- mcsamplerNFlow was missing from the list the docstring calls exhaustive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tical to base Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…opt-in Prototype opt-in precession-informed RF features with low-mass Asimov policy
- test_calmarg_rift_source: stop at the hlmoft call and compare modes, so the test no longer reaches bilby's utils.nfft (bilby is not installed in CI). - _dry_run_argv: only \" is escaped in the job ad; keep backslashes. Add a backslash case and a $(macro)-inside-a-value case; fail on unsubstituted macros. - On a failed DAG build, print every error block, not only the last 4000 characters (child output is interleaved out of order). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reference bilby ini names calibration envelopes under a CIT home directory, which CEPP copies at build time, so every DAG build failed on the GitHub runner. Write a copy of the bilby ini that points at empty stub files in tmp_path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
run_lisa_known_sky_surface built its CEPP command without --ile-request-disk, --cip-request-disk or --general-request-disk, so every LISA .sub file got CEPP's default request_disk = 10M. Forward the three --internal-*-request-disk options as the normal path does, and pin each .sub file's request_disk in the LISA contract test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
calibration_reweighting.py now starts from fd_alignment_postevent_time=2, rift_source adds fd_standoff_factor=0.9 on the lalsimutils.hlmoft route unless set, and pseudo_pipe nests --internal-waveform-extra-lalsuite-args under 'extra_waveform_args'. These are the options ILE passes (analyze_event; factored_likelihood.internal_hlm_generator), with the user's --internal-waveform-extra-kwargs winning as in ILE. test_calmarg_ile_waveform_alignment.py builds real DAGs and compares the generator options calmarg and ILE reach, and their XPHM/XO4a modes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…art consolidate working Resolve the core-unit floor conflict, then fix review findings: - util_ILEdagPostprocess.sh: with no CME*.dat shards, write an empty composite and exit 0 as before. Under --first-iteration-jumpstart the first consolidate node has no ILE parents and BasicIteration omits its nonempty-composite POST check, so failing there stopped the DAG at its first node. Resolve each cleaner only in the branch that runs it, and fail (instead of copying) when the legacy cleaner yields no usable rows. - util_NRdagPostprocess.sh: take the relabeler's status from PIPESTATUS, so an empty grep match is reported as an empty index, not a relabel failure. - Tests: cover jumpstart, the hyperpipeline branch, and NR cleaner and relabel failures (13 tests); assert the join's exit status in test_cleanile_intrinsic_precision.py now that it no longer always exits 0. - Core-unit floors 671/658. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…disk-requests pseudo_pipe: forward disk requests on the LISA known-sky path
… ast.unparse Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
util_NRRelabelILE.py prints the best-match lnL first. For a real signal it is positive, so the `grep '^-1*'` filter normally keeps nothing, and no builder reads .indexed. Keep the relabeler exit-status check; drop the nonempty check. Record the core-unit gate measurement. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…process-path-o4d Fail closed in DAG postprocessing
…-args-quoting calmarg: transport waveform-kwargs dictionaries with typed values
…aveform-alignment-o4d calmarg: build the same waveform as ILE (alignment kwargs, lalsuite args)
…erpolator-repair Resolved as on claude/trial-382-on-387 (5d3037b): one _CIP_FIT_METHODS list with gp-matern; fitting-method check only in _validate_cip_fit_method; both asimov validators in before_config and build_dag; helper gp-matern rewrite before the RF transverse block, and gp-matern with RF transverse auto/physics3 is a parser.error; CIP and rift.ini keep both option sets. test-core-units.sh: both sides' floor increments added (691/677, 14 skips). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…input masses With source_redshift=z, the vectorized mu1/mu2 blocks used the source-frame mc. Now the low-level mass inputs (mc, mc_ecc, m1, m2, mtot) are multiplied by 1+z once, on a copy, at entry, and the end-of-function rescale is gone. The per-row fallback rescales P.m1, P.m2 only when no input was rescaled. Test: vectorized output vs the per-row extract_param path at z=0 and 0.3, including the in-plane ring coordinates; registered in core-units (686/673). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s it where importable Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…polator-repair Repair GPU GP interpolation and add experimental AV stopping choice
…urce-redshift-detector-frame convert_waveform_coordinates: apply source_redshift to input masses (mu1/mu2 frame)
oshaughnessy-junior
deployed
to
private-review-dispatch-rift-upstream
October 5, 2026 12:08 — with
GitHub Actions
Active
…-resolution Resolve upstream PR 208 release-history conflict
oshaughnessy-junior
deployed
to
private-review-dispatch-rift-upstream
October 5, 2026 12:17 — with
GitHub Actions
Active
oshaughn
deployed
to
private-review-dispatch-rift-upstream
October 5, 2026 13:14 — with
GitHub Actions
Active
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bring
oshaughnessy-junior:rift_O4dintooshaughn:rift_O4d, following the cross-fork delivery convention of #188, #191, and #193.This is the next catch-up after #193: evidence and prior corrections, sampler density/backend repairs, CIP/EOS and hyperpipe improvements, Asimov compatibility, workflow/container fixes, calibration waveform consistency, and expanded validation.
oshaughn/rift_O4d@f50713beadcd88adb9ab3844cd1dadb4edec93ed.oshaughnessy-junior/rift_O4d@7e2984a03b2f4b48ea9f508ac2c9d19463ef5180(after conflict resolution via junior #398).rift_O4d.Individual PRs by subsystem
All
junior #Nlinks below refer tooshaughnessy-junior/research-projects-RIT, not upstream PR numbers. Each entry describes the constituent PR; later entries can refine or supersede earlier implementations.Evidence, priors, and posterior products
Sampler densities, portfolio contracts, and backend repairs
CIP, EOS, hyperpipe, and coordinate conversion
JAX priors and interpolation experiments
Workflow, Asimov, containers, and resource requests
Calibration waveform consistency
Validation and CI
Agent documentation
Scientific and operational behavior changes
ln(2), orln(8)with--internal-rotate-phase; posterior shapes are unchanged. Oldall.net/composite rows and reused CIP fits carry the former offset and must be reconciled before combining them with new results.chi_maxdiffers from one. Add ASIMOV 0.7 compatibility and PESummary assets #176 adds a Bilby-compatible fixed-PSD noise-evidence sidecar, with its convention documented in that PR.+500. The constituent PR calls for RO approval of this behavior change.Integration and release notes
Junior PR #398 merged upstream
rift_O4d@f50713beainto the fork, clearing the only conflict, inCHANGES.rst. It preserves upstream's rc0–rc5 release history verbatim and keeps the additional fork JAX terminal-output/EOS notes in an unreleased entry.setup.pyadopts upstream's existing0.0.18.0rc5version; no additional version bump was introduced. The resolution delta against junior is limited to these two metadata files, with no runtime source changes.Validation of the resolution: exact preservation checks for upstream history/version, fork-only notes, and older release history; no conflict markers or unmerged index entries;
git diff --check; Python AST parsing ofsetup.py; both prior tips are ancestors of the resolution. Delivered from an isolated temporary checkout using junior HTTPS credentials; the human checkout was unchanged.GitHub now reports
MERGEABLE, with zero upstream-only commits. Its current merge-state status isUNSTABLEwhile checks on the updated head run. This resolves branch conflicts; upstream #208 remains open for review and validation.Validation
At the pre-resolution head
1997e6c3c, all 34 GitHub check runs completed successfully in fork CI run 37296026318. These include core units, integration, three JAX shards, calmarg, LISA, slow rotation, Asimov 0.5/0.7/0.8, Rimsky, dependency/import checks, install checks on Python 3.9–3.13, container canaries, docs, and CI roster/accounting checks.Fresh CI on the updated head is still running; the earlier 34 green checks apply to the pre-resolution head. No additional aggregate or production/GPU science suite was rerun locally for the metadata-only resolution. Feature-specific numerical, mutation, GPU, and production evidence remains linked from the individual PRs above; expensive posterior shape-recovery validation is not newly established by this roll-up description.