Skip to content

fix(toolkit-lib): restore the replacement guard for Express Mode deployments - #1969

Draft
sanjanaravikumar-az wants to merge 12 commits into
mainfrom
fix/express-replacement-guard
Draft

sanjanaravikumar-az wants to merge 12 commits into
mainfrom
fix/express-replacement-guard

Conversation

@sanjanaravikumar-az

@sanjanaravikumar-az sanjanaravikumar-az commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #1931
Express Mode disables rollback by default, and CloudFormation refuses to replace a resource while rollback is disabled. The CLI used to catch that before submitting anything and offer to re-run with rollback enabled. The check for express deploys was removed, so since 2.1133.0 the CLI submits deployments CloudFormation is guaranteed to reject and the stack is left in UPDATE_FAILED.

CloudFormation accepts the change set and accepts the execute call. It only refuses once execution reaches the resource, which is why an operation has already started and the stack ends up stuck rather than erroring immediately. In the single-resource stack I tested, express rolled the rejected resource back on its own (UPDATE_ROLLBACK_COMPLETE "Rollback succeeded for the failed resources.") and no orphaned resource was left behind — but that is only what one resource does. A multi-resource update can land other changes before the rejection, and with rollback disabled those stay, so the guard is not merely cosmetic.

A way to get out of this state is to have a deployment that doesn't change anything:

  1. Revert the change so the app matches the last configuration that deployed successfully.
  2. Run cdk deploy --express --method=direct which replays that configuration as a no-op and the stack is at UPDATE_COMPLETE. --method=direct bypasses the changeset check and goes to UpdateStack directly.
  3. Re-apply the changes and deploy it with cdk deploy --express --rollback.

Changes:

  • Restore the replacement check, and point the user at cdk deploy --express --rollback.

  • Detect the rejection after the fact as well, from monitorDeployment, which both the change set path and --method=direct run through. It is not a --method=direct-only hook: --method=direct has no change set to inspect, and separately a change set can report Replacement: "Conditional", which the up-front gate does not count as a replacement (#1971) — this is the only thing that fires if CloudFormation later resolves it to a real replacement. The original CloudFormation error still propagates.

  • Order the replacement guidance ahead of the generic failure reporting in monitorDeployment. This changes the order of what every failing deployment prints, standard mode included, not just express.

  • From an already-failed stack, fail with a terminal error carrying the unwedge steps.

  • When executing an existing change set, read the rollback policy off the change set rather than the current command, because a change set freezes that policy at creation. Where the persisted policy disagrees with the current flags but there is no replacement, this now warns and proceeds rather than refusing — CloudFormation accepts those executions, so refusing them would break deployments unrelated to (cli): #1745 removed the express-mode replacement guard, so --express deploys replacements with rollback disabled #1931. A replacement under a frozen rollback-disabled policy is terminal: the change set has to be recreated, because retrying it with rollback enabled re-executes the same frozen change set.

  • Gate replacements inside nested stacks. The root change set can be replacement-free while a child change set carries the replacement; that bypassed the guard entirely and executed.

  • Stop sending DisableRollback on ExecuteChangeSet when the change set already pins the policy. commonExecuteOptions() sends DisableRollback: true for an explicit --no-rollback, which conflicts with a change set created with --rollback and made CloudFormation reject the call outright.

  • Forward express to toolkit-lib on the execute-change-set path. cdk-toolkit.ts delegates that method straight to toolkit.deploy() and passed rollback but not express, so the execute path always took the non-express branch. On main that means --express --method=execute-change-set can wedge a stack today — this is a pre-existing bug, not a refusal this PR introduces. express has to travel with rollback because Express Mode flips what a missing rollback means, and without it toolkit-lib read plain --express as rollback-enabled.

  • Recovery guidance depends on the stack's actual state, for example a failed initial create has no previous configuration to replay.

  • Restored the regression test for the express replacement check that was removed alongside the guard.

  • Fixes a diagnostic that reported rollback as enabled under express and advised re-running with --no-rollback, and makes the cdk deploy confirmation prompt use the same wording on both CLI paths, so express users are no longer told a flag they never passed is the problem.

Notes:

  • There are three distinct rollback predicates in deploy-stack.ts and they are deliberately not unified, because they answer different questions: what the request asks for, what an already-created change set will actually do, and the raw flag default used only for standard-mode paused-state routing. They are documented as such rather than collapsed.

  • Passing DisableRollback at execute time is not a workaround for this. It is a consistency assertion, not an override: a matching value is accepted, and a conflicting one fails synchronously with ValidationError: DisableRollback specified on ExecuteChangeSet conflicts with the value DisableRollback the ChangeSet was created with. That is why the CLI cannot honour a late --rollback by passing it through, and why refusing up front produces a better message than the service error would.

  • The actual rule is that the rollback policy is fixed at change set creation: a change set created with --deployment-config '{"Mode":"EXPRESS","DisableRollback":false}' executes the same replacement successfully (exit 0, task definition revision :1 → :2).

  • Guard removal is a single unit and is documented in the code: the replacement branch, the mismatch report, the after-the-fact detection hook, the ReplacementRequiresRollback/ReplacedResource payload types and CDK_TOOLKIT_W5903 all come out together. A server-side CloudFormation fix was reported to us second-hand for around 2026-11-15; that is unconfirmed and nothing should be removed before the restriction is observed to be gone.

Verified against CloudFormation on a live ECS stack in us-east-1: the flag combination is accepted silently and persists {"Mode":"EXPRESS","DisableRollback":true}; executing it with --rollback wedges the stack in UPDATE_FAILED with Replacement type updates not supported on stack with disable-rollback; omitting --express on the execute wedges it identically; the guard then submits nothing (stack event count and LastUpdatedTime unchanged, change set left intact); the documented unwedge works; and retrying with rollback from a failed stack fails as described.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache-2.0 license.

…oyments

CloudFormation rejects replacement-type updates while rollback is disabled.
Express Mode disables rollback unless --rollback is passed, so `cdk deploy
--express` submits replacements CloudFormation refuses, and the stack is left in
UPDATE_FAILED with no rollback available. #1745 removed the express half of the
guard condition (and relaxed the test covering it), #1785 restructured what was
left into `if (!this.options.express)`.

Derive the condition once in `rollbackDisabled()` and apply the guard in both
modes, returning `replacement-requires-rollback` so the existing
confirm-and-retry-with-rollback prompt engages. Fix the same condition in the
failure diagnostic, which reported rollback as enabled under --express and
advised users to re-run with --no-rollback.

`--method=direct` is deliberately not refused up front, because redeploying the
previous configuration that way is how a stuck stack is unwedged; instead the
rejection is detected after the fact and the user is routed to `--express
--rollback`. That routing reads the activity monitor's errors, so the monitor is
now flushed before they are read.

Fixes #1931
@github-actions

Copy link
Copy Markdown
Contributor

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@aws-cdk-automation
aws-cdk-automation requested a review from a team September 18, 2026 05:59
… gated

Move three rationales into the code, where they are durable:

- --express --method=direct is deliberately not refused up front, because
  replaying the previous configuration that way is the only exit from a stack
  already stranded in UPDATE_FAILED.
- Replacement: 'Conditional' is deliberately excluded, because CDKMetadata
  reports it on essentially every CDK deployment.
- The CloudFormation reason match should be replaced with a structured
  discriminator if one ever becomes available.
sanjrkmr and others added 9 commits September 18, 2026 20:34
…terminally when already wedged

Two defects in the replacement guard.

The direct-path routing was gated on rollbackDisabled() alone, which is also true
for a standard-mode `--no-rollback` deployment. Since the guidance names Express
Mode flags, a wedged standard-mode user was told to switch to Express Mode, which
is sticky and gives up `cdk rollback` - a strictly worse position for a problem a
plain `cdk deploy` fixes. The notify is now gated on `express`, matching the
change-set arm.

On express with a replacement against an already-failed stack, the guard printed
that `--express --rollback` cannot update the stack and then returned
`replacement-requires-rollback`, so the toolkit offered exactly that deployment.
The confirmation defaults to yes, so a non-interactive caller ran it and hit
`ValidationError: non-terminal [UPDATE_FAILED]`. That case now throws
`ReplacementRequiresUnwedge` carrying the unwedge steps, so it is reported once
and CI cannot auto-run a deployment that is known to fail.

Also: `ReplacedResource.logicalId` is optional and both fabricated fallbacks
('<unknown>' and the stack name) are gone, so `replacements` can be genuinely
empty as documented; `shouldDisableRollback` renamed to
`shouldSendDisableRollbackFlag`; and the comment claiming express cannot send a
top-level `DisableRollback` is corrected (`--express --no-rollback` does send it).
… current invocation

A change set carries its own rollback policy. CloudFormation persists
DeploymentConfig on CreateChangeSet, returns it from DescribeChangeSet, and
ExecuteChangeSet has no DeploymentConfig field, so it cannot override it. The
guard read the current invocation's --express/--rollback flags instead, which is
only correct when one invocation both creates and executes the change set.

Split across invocations - `--method=change-set --no-execute` then
`--method=execute-change-set`, or an explicit execute-change-set retry - the
second command's flags decided nothing. Passing --rollback, or simply omitting
--express, made the guard conclude rollback was enabled and execute a
still-rollback-disabled change set containing a replacement, putting the
prohibited combination back in front of CloudFormation and reaching #1931 again.

Rollback state now comes from the persisted DeploymentConfig when it pins the
answer (Express only; standard mode decides at execute time via the
DisableRollback flag, so current options stay authoritative there). A requested
policy that conflicts with the persisted one is refused with
ChangeSetRollbackPolicyMismatch telling the user to create a new change set,
rather than silently doing the opposite of what was asked.

The fake could not express any of this because its DescribeChangeSet dropped
DeploymentConfig; it now returns it, and createChangeSetSync accepts one.

Also restricts the replay recovery guidance to UPDATE_FAILED. isRollbackable also
covers CREATE_FAILED and UPDATE_ROLLBACK_FAILED, and a stack that never deployed
successfully has no previous configuration to replay, so those states get their
own guidance instead of impossible instructions.
`--express` decides what a missing `--rollback` means: Express Mode disables
rollback by default, standard mode enables it. The `execute-change-set`
delegation in `CdkToolkit.deploy` enumerates its options explicitly and dropped
`express`, so toolkit-lib read plain `--express` as "rollback enabled" and the
change-set policy guard refused Express change sets this same CLI had just
created with rollback disabled:

    cdk deploy --express --method=prepare-change-set   # persists DisableRollback: true
    cdk deploy --express --method=execute-change-set   # ChangeSetRollbackPolicyMismatch

Forward `express` alongside `rollback` so the policy is derived identically on
both paths. The guard stays unscoped to `--express`: a persisted EXPRESS policy
still governs regardless of this invocation's flags, so omitting `--express` on
execute is still refused.

Also derive the effective policy for the replacement guard from what will
actually be sent. Only Express records a rollback choice on the change set;
anything else is governed by the `DisableRollback` on `ExecuteChangeSet`, which
is only sent for an explicit `--no-rollback`. Without this, forwarding `express`
would turn a standard `prepare` followed by `execute --express` into a false
refusal.

Replace the "undocumented" note on whether execute-time `DisableRollback` can
override a persisted EXPRESS policy with the verified behaviour: it is a
consistency assertion, not an override. A matching value is accepted; a
conflicting one fails synchronously with `ValidationError: DisableRollback
specified on ExecuteChangeSet conflicts with the value DisableRollback the
ChangeSet was created with.`
`cdk deploy --express --rollback` cannot update a stack that is already failed -
CloudFormation answers "This stack is currently in a non-terminal [UPDATE_FAILED]
state" - so the guidance has to lead with the state the stack is in, then the steps
that return it to a terminal state, and only then suggest deploying with rollback.
The existing tests asserted only that those strings were present, so a reordering
that led with `--rollback` would have shipped green.

Compare offsets instead, at both sites that render the `replay` variant. The
assertion deliberately anchors on the suggestion phrasing rather than on
`indexOf('--rollback')`: the first occurrence of that flag is the clause saying it
cannot update a failed stack, which legitimately precedes the unwedge steps, so a
naive "unwedge before any --rollback mention" check would fail on correct output.

Also pin the `recreate` variant, whose wrongness is hardest to spot by hand: a stack
that never completed a deployment has no configuration to replay, so that guidance
must not contain replay instructions at all.
…nt guard

The confirm-and-retry prompt is not the only recovery path: it is what a healthy
stack gets, because the guard returns `replacement-requires-rollback` there. An
already-failed stack throws `ReplacementRequiresUnwedge` instead, since the retry
cannot succeed from that state.

Scope the fake's rollback note to replacements. Rollback being disabled does not by
itself strand a stack - an ordinary express deployment with no replacement completes
normally. It is a replacement submitted while rollback is disabled that CloudFormation
rejects during execution.

Say `ExecuteChangeSet` "cannot change" the persisted `DeploymentConfig` rather than
"cannot override" it: a `DisableRollback` matching the change set is accepted, and
only a conflicting one is rejected.
…change set executions

Address review on the express replacement guard:

- Gate replacements found in nested stacks' change sets, which previously
  bypassed the guard entirely because the root change set is replacement-free.
- Stop refusing a persisted/requested rollback policy mismatch when there is no
  replacement; CloudFormation accepts those, so warn and proceed instead.
- Make a replacement under a frozen rollback-disabled policy terminal, since
  retrying re-executes the same change set and cannot succeed.
- Read the persisted DeploymentConfig before paused-state routing, so an express
  change set is not sent down the standard rollback-first path.
- Emit W5903 before throwing, so 'cdk deploy --watch' still shows the guidance.
- Omit DisableRollback on ExecuteChangeSet when the change set pins the policy;
  sending it for an explicit --no-rollback conflicted with a change set created
  with --rollback and made CloudFormation reject the call.
- Fail closed when DeploymentConfig is absent on an express invocation.
- Use the express-aware confirmation wording on both CLI deploy paths.
- Document the three distinct rollback predicates, the shared detection hook,
  the unit-of-removal contract, and that the 2026-11-15 CloudFormation fix is
  second-hand and unconfirmed.
…ge steps

Mutation testing found a gap: prepending an actionable
'Deploy it with cdk deploy --express --rollback' line ahead of the whole
explanation passed every existing assertion, because the unwedge steps still
followed 'in a failed state' and the closing suggestion still came last. The
helper deliberately avoided indexOf('--rollback') to allow the legitimate
'cannot update' clause, and that carve-out was the blind spot.

Pin it explicitly: the first mention of the rollback command must be the
'cannot update' clause. The mutation now fails 2 tests instead of 0.

This branch was successfully deployed

3 active (1 outdated) deployments
run-tests — 2200e3d9 Deployed Sep 29, 2026 by sanjanaravikumar-az via integ_cli (cli-integ-tests, 24.19, 5) #7028
no-approval — 2200e3d9 Deployed Sep 29, 2026 by sanjanaravikumar-az via prepare #7028
automation — 9a9caf7f Deployed Sep 18, 2026 by sanjanaravikumar-az via Triage Pull Requests #2102
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

(cli): #1745 removed the express-mode replacement guard, so --express deploys replacements with rollback disabled

3 participants