Function
Output sequence
The default output sequence of AlphaImpute2 is by the ascending sequence of individual ID, unlike AlphaPeel's default, which preserves the input files' sequence; this requires highlighting in some places
Runtime
Numba compilation
The compilation is taking a long time; the profiled result gives the following time taken for the compilation:
name ncall tsub ttot tavg
..43 CPUDispatcher._compile_for_args 42 0.000755 87.97237 2.094580
The compiled code is saved in RAM, so each new run submitted to the work node would lead to a new round of compilation. To save the compilation overhead and get more accurate runtime results, Ahead-of-Time compilation may be worth considering.
RNG
The random number generation for the particle phasing is taking a long time:
name ncall tsub ttot tavg
..lePhasing.py:118 get_random_values 110.. 117.4144 117.6782 0.001070
AlphaImpute2 has a unique random number generator for each individual; each has a unique seed that is dependent on the individual's ID via a complicated generation process (hashing + multiply with a very big number). In the real scenario, the random number generator is switched between different states as the individuals to phase is changed frequently, and the re-setup may take a long time.
While I didn't fully get why we need a completely independent random number generator for each individual (I think it should be safe enough if we have it for each thread). Suppose we require the very independent seed values. The current seed value generator is hard-coded in get_random_seed. I found an alternative by NumPy: SeedSequence.
While the time on average for the generation, I estimated the size of the array for each call of the random number generation, and submitted the following script to the work node with the same memory request
python -m timeit -s "import numpy as np; rng = np.random.default_rng()" "rng.random(25000)"
and I get
5000 loops, best of 5: 79.2 usec per loop
While the tavg above is 13 times slower than this. I would suspect a long time is taken just to switch the random number generator from one to another.
Repetitive calculation
There are places such as the following:
new[0] = e2 * tmp[3] + e1e * (tmp[1] + tmp[2]) + e2i * tmp[0]
new[1] = e2 * tmp[2] + e1e * (tmp[0] + tmp[3]) + e2i * tmp[1]
new[2] = e2 * tmp[1] + e1e * (tmp[0] + tmp[3]) + e2i * tmp[2]
new[3] = e2 * tmp[0] + e1e * (tmp[1] + tmp[2]) + e2i * tmp[3]
which the calculation of tmp[1] + tmp[2] and (tmp[0] + tmp[3]) can be put in advance and reduce the repetitive calculation. Sometimes Numba may have already done the optimisation during the compilation (for this particular case, inspect_types() tells me it did twice the same additions, so it still needs optimisation by hand).
Memory safety
AlphaImpute2 has a much higher rate of page-faults than AlphaPeel (about 100 times more per second). Considering the unknown cause of the segmentation fault when calling AlphaImpute2, it is very likely that some unsafe manipulations are in the code.
Global variables
One possible cause might be the global variables in Heuristic_Peeling.py, there are some discussions here: https://www.quora.com/Why-are-global-variables-generally-frowned-upon-What-are-some-of-the-risks-associated-with-using-them
This article also mentioned the possible issues of global variables with multi-threading.
We may consider some other alternatives to the global variables.
Function
Output sequence
The default output sequence of AlphaImpute2 is by the ascending sequence of individual ID, unlike AlphaPeel's default, which preserves the input files' sequence; this requires highlighting in some places
Runtime
Numba compilation
The compilation is taking a long time; the profiled result gives the following time taken for the compilation:
The compiled code is saved in RAM, so each new run submitted to the work node would lead to a new round of compilation. To save the compilation overhead and get more accurate runtime results, Ahead-of-Time compilation may be worth considering.
RNG
The random number generation for the particle phasing is taking a long time:
AlphaImpute2 has a unique random number generator for each individual; each has a unique seed that is dependent on the individual's ID via a complicated generation process (hashing + multiply with a very big number). In the real scenario, the random number generator is switched between different states as the individuals to phase is changed frequently, and the re-setup may take a long time.
While I didn't fully get why we need a completely independent random number generator for each individual (I think it should be safe enough if we have it for each thread). Suppose we require the very independent seed values. The current seed value generator is hard-coded in
get_random_seed. I found an alternative by NumPy:SeedSequence.While the time on average for the generation, I estimated the size of the array for each call of the random number generation, and submitted the following script to the work node with the same memory request
and I get
While the
tavgabove is 13 times slower than this. I would suspect a long time is taken just to switch the random number generator from one to another.Repetitive calculation
There are places such as the following:
which the calculation of
tmp[1] + tmp[2]and(tmp[0] + tmp[3])can be put in advance and reduce the repetitive calculation. Sometimes Numba may have already done the optimisation during the compilation (for this particular case,inspect_types()tells me it did twice the same additions, so it still needs optimisation by hand).Memory safety
AlphaImpute2 has a much higher rate of page-faults than AlphaPeel (about 100 times more per second). Considering the unknown cause of the segmentation fault when calling AlphaImpute2, it is very likely that some unsafe manipulations are in the code.
Global variables
One possible cause might be the global variables in
Heuristic_Peeling.py, there are some discussions here: https://www.quora.com/Why-are-global-variables-generally-frowned-upon-What-are-some-of-the-risks-associated-with-using-themThis article also mentioned the possible issues of global variables with multi-threading.
We may consider some other alternatives to the global variables.