Skip to content

Allow configurability of liveness and readiness probe. #1732

Description

Context: Wing is investigating why our CRDB pods seem to die on a regular basis with OOM.

We fixed the resource-limit issue - but pods still died.

We moved to a higher memory machine and pods still died. However: the last pod death is no longer about memory limit exceeded but rather the LivenessProbe which has a very strict timeout of 1 second.

It appears that, liveness probe for CRDB is not recommended. After all, CRDB is repeatedly checking group membership own its own. During high load (regardless of client-caused or database-internals-caused), a liveness probe can easily fail.

We should remove the LivenessProbe and allow the users to configure the readiness probe as well.

While we still don't know why there is an hourly memory spike in one of the 3 pods (would be great if we could just ask Cockroachdb team), we do now know that liveness probe is not recommended. If the hourly memory spike is a normal thing somehow: then liveness probe removal is the next thing we can do to help our pods. Online searches show other people facing issues with liveness probe as well cockroachdb/cockroach#44832

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1High prioritydeploymentRelated to deploying a DSS instance rather than application logic or behaviordssRelating to one of the DSS implementations

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions