Describe the bug
The Test replication step of .github/workflows/build.yml stands up three servers on a numbering scheme of
N389 / N636 / N4444 / N8989. At N=3 that scheme crosses a boundary it was not written against:
| server |
ldap |
ldaps |
admin connector |
replication |
| OpenDJ-1 |
1389 |
1636 |
4444 |
8989 |
| OpenDJ-2 |
2389 |
2636 |
24444 |
28989 |
| OpenDJ-3 |
3389 |
3636 |
34444 |
38989 |
Of the twelve numbers, exactly two sit above 32768, and both belong to the third server. The Linux default
net.ipv4.ip_local_port_range is 32768 60999, so 34444 and 38989 are inside the range the kernel hands out
as local source ports for outgoing connections. The step runs after a full Maven build, a Cassandra step and a
Postgres step, so there is no shortage of outbound sockets on the runner; when one of them holds 34444 at the
wrong moment, the third server cannot have it.
Nothing in the workflow binds these ports twice — the assignment is internally consistent. The collision is with
whatever else on the machine happens to be dialling out.
Seen in CI
Two failures on 2026-09-11, both at Setup OpenDJ-3 with replication, both on branches that touch only the JDBC
backend, and on legs whose siblings were green in the same run.
run 34564104369, build-maven (ubuntu-latest, 21) — caught by setup's own port check, 1.3 s after the step echoed its banner:
Setup OpenDJ-3 with replication
ERROR: Unable to bind to port 34444. This port may already be in use, or you
may not have permission to bind to it
##[error]Process completed with exit code 2.
run 34565184252, build-maven (ubuntu-latest, 17) — past the port check, configured, base entry created, and then unable to start:
Setup OpenDJ-3 with replication
Configuring Directory Server ..... Done.
Configuring Certificates ..... Done.
Creating Base Entry dc=example,dc=com ..... Done.
Starting Directory Server ......
Error Starting Directory Server. Error code: 1.
##[error]Process completed with exit code 7.
The second one does not name a port: the detail log it points at is unreadable, which is #1030. The two together
are what the race looks like from either side of setup's pre-flight check — free when asked, taken when bound.
Across the 60 most recent failed workflow runs, these are the only two whose Test replication step failed. Rare,
but it costs a full leg each time and lands on branches that have nothing to do with replication.
Suggested fix
Renumber the administrative ports of servers 2 and 3 into the block next to server 1, well clear of the ephemeral
range:
|
before |
after |
| OpenDJ-2 admin connector |
24444 |
4445 |
| OpenDJ-2 replication |
28989 |
8990 |
| OpenDJ-3 admin connector |
34444 |
4446 |
| OpenDJ-3 replication |
38989 |
8991 |
Server 2's current ports are below 32768 and are not at risk today; they are in the change so that the scheme stops
being positional. 4444/4445/4446 and 8989/8990/8991 are contiguous, obviously bounded, and a fourth server
would extend them without anyone having to re-check where the ephemeral range starts — which is the trap the
present N4444 scheme sets.
None of 4445, 4446, 8990, 8991 occurs anywhere in the repository today, and none collides with the other
ports the same job uses (1389, 1636, 4444, 5432 for Postgres, 9042 for Cassandra). The LDAP and LDAPS
ports (1389/2389/3389, 1636/2636/3636) are already below the boundary and do not change.
Seven lines in .github/workflows/build.yml carry the four numbers:
:309 opendj2/setup --adminConnectorPort 24444
:314 dsreplication enable --port2 24444 --replicationPort2 28989
:318 dsreplication initialize --portDestination 24444
:325 opendj3/setup --adminConnectorPort 34444
:329 dsreplication enable --port1 24444 --replicationPort1 28989
:330 dsreplication enable --port2 34444 --replicationPort2 38989
:334 dsreplication initialize --portSource 24444 --portDestination 34444
Note
The net.ipv4.ip_local_port_range value is the documented Linux default rather than something read off a GitHub
runner — the workflow does not print it. It is the mechanism that fits what is observed: only the third server is
ever affected, only on Linux, intermittently, and at whichever of the two binds happens to lose. Raising the range
with sysctl in the step would also work, but it needs root and changes machine-wide behaviour to fix four
numbers.
Describe the bug
The
Test replicationstep of.github/workflows/build.ymlstands up three servers on a numbering scheme ofN389 / N636 / N4444 / N8989. AtN=3that scheme crosses a boundary it was not written against:Of the twelve numbers, exactly two sit above 32768, and both belong to the third server. The Linux default
net.ipv4.ip_local_port_rangeis32768 60999, so34444and38989are inside the range the kernel hands outas local source ports for outgoing connections. The step runs after a full Maven build, a Cassandra step and a
Postgres step, so there is no shortage of outbound sockets on the runner; when one of them holds
34444at thewrong moment, the third server cannot have it.
Nothing in the workflow binds these ports twice — the assignment is internally consistent. The collision is with
whatever else on the machine happens to be dialling out.
Seen in CI
Two failures on 2026-09-11, both at
Setup OpenDJ-3 with replication, both on branches that touch only the JDBCbackend, and on legs whose siblings were green in the same run.
run 34564104369,
build-maven (ubuntu-latest, 21)— caught by setup's own port check, 1.3 s after the step echoed its banner:run 34565184252,
build-maven (ubuntu-latest, 17)— past the port check, configured, base entry created, and then unable to start:The second one does not name a port: the detail log it points at is unreadable, which is #1030. The two together
are what the race looks like from either side of setup's pre-flight check — free when asked, taken when bound.
Across the 60 most recent failed workflow runs, these are the only two whose
Test replicationstep failed. Rare,but it costs a full leg each time and lands on branches that have nothing to do with replication.
Suggested fix
Renumber the administrative ports of servers 2 and 3 into the block next to server 1, well clear of the ephemeral
range:
Server 2's current ports are below 32768 and are not at risk today; they are in the change so that the scheme stops
being positional.
4444/4445/4446and8989/8990/8991are contiguous, obviously bounded, and a fourth serverwould extend them without anyone having to re-check where the ephemeral range starts — which is the trap the
present
N4444scheme sets.None of
4445,4446,8990,8991occurs anywhere in the repository today, and none collides with the otherports the same job uses (
1389,1636,4444,5432for Postgres,9042for Cassandra). The LDAP and LDAPSports (
1389/2389/3389,1636/2636/3636) are already below the boundary and do not change.Seven lines in
.github/workflows/build.ymlcarry the four numbers:Note
The
net.ipv4.ip_local_port_rangevalue is the documented Linux default rather than something read off a GitHubrunner — the workflow does not print it. It is the mechanism that fits what is observed: only the third server is
ever affected, only on Linux, intermittently, and at whichever of the two binds happens to lose. Raising the range
with
sysctlin the step would also work, but it needs root and changes machine-wide behaviour to fix fournumbers.