Skip to content

CI: the third replication server is given ports inside the ephemeral range, so its setup fails at random #1031

Description

@vharseko

Describe the bug

The Test replication step of .github/workflows/build.yml stands up three servers on a numbering scheme of
N389 / N636 / N4444 / N8989. At N=3 that scheme crosses a boundary it was not written against:

server ldap ldaps admin connector replication
OpenDJ-1 1389 1636 4444 8989
OpenDJ-2 2389 2636 24444 28989
OpenDJ-3 3389 3636 34444 38989

Of the twelve numbers, exactly two sit above 32768, and both belong to the third server. The Linux default
net.ipv4.ip_local_port_range is 32768 60999, so 34444 and 38989 are inside the range the kernel hands out
as local source ports for outgoing connections. The step runs after a full Maven build, a Cassandra step and a
Postgres step, so there is no shortage of outbound sockets on the runner; when one of them holds 34444 at the
wrong moment, the third server cannot have it.

Nothing in the workflow binds these ports twice — the assignment is internally consistent. The collision is with
whatever else on the machine happens to be dialling out.

Seen in CI

Two failures on 2026-09-11, both at Setup OpenDJ-3 with replication, both on branches that touch only the JDBC
backend, and on legs whose siblings were green in the same run.

run 34564104369, build-maven (ubuntu-latest, 21) — caught by setup's own port check, 1.3 s after the step echoed its banner:

Setup OpenDJ-3 with replication
ERROR:  Unable to bind to port 34444.  This port may already be in use, or you
may not have permission to bind to it
##[error]Process completed with exit code 2.

run 34565184252, build-maven (ubuntu-latest, 17) — past the port check, configured, base entry created, and then unable to start:

Setup OpenDJ-3 with replication
Configuring Directory Server ..... Done.
Configuring Certificates ..... Done.
Creating Base Entry dc=example,dc=com ..... Done.
Starting Directory Server ......
Error Starting Directory Server.  Error code: 1.
##[error]Process completed with exit code 7.

The second one does not name a port: the detail log it points at is unreadable, which is #1030. The two together
are what the race looks like from either side of setup's pre-flight check — free when asked, taken when bound.

Across the 60 most recent failed workflow runs, these are the only two whose Test replication step failed. Rare,
but it costs a full leg each time and lands on branches that have nothing to do with replication.

Suggested fix

Renumber the administrative ports of servers 2 and 3 into the block next to server 1, well clear of the ephemeral
range:

before after
OpenDJ-2 admin connector 24444 4445
OpenDJ-2 replication 28989 8990
OpenDJ-3 admin connector 34444 4446
OpenDJ-3 replication 38989 8991

Server 2's current ports are below 32768 and are not at risk today; they are in the change so that the scheme stops
being positional. 4444/4445/4446 and 8989/8990/8991 are contiguous, obviously bounded, and a fourth server
would extend them without anyone having to re-check where the ephemeral range starts — which is the trap the
present N4444 scheme sets.

None of 4445, 4446, 8990, 8991 occurs anywhere in the repository today, and none collides with the other
ports the same job uses (1389, 1636, 4444, 5432 for Postgres, 9042 for Cassandra). The LDAP and LDAPS
ports (1389/2389/3389, 1636/2636/3636) are already below the boundary and do not change.

Seven lines in .github/workflows/build.yml carry the four numbers:

:309  opendj2/setup      --adminConnectorPort 24444
:314  dsreplication enable  --port2 24444 --replicationPort2 28989
:318  dsreplication initialize  --portDestination 24444
:325  opendj3/setup      --adminConnectorPort 34444
:329  dsreplication enable  --port1 24444 --replicationPort1 28989
:330  dsreplication enable  --port2 34444 --replicationPort2 38989
:334  dsreplication initialize  --portSource 24444 --portDestination 34444

Note

The net.ipv4.ip_local_port_range value is the documented Linux default rather than something read off a GitHub
runner — the workflow does not print it. It is the mechanism that fits what is observed: only the third server is
ever affected, only on Linux, intermittently, and at whichever of the two binds happens to lose. Raising the range
with sysctl in the step would also work, but it needs root and changes machine-wide behaviour to fix four
numbers.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions