Skip to content

fix(upgrade): don't exit silently when .env lacks AGENT_SERVER_IMAGE - #549

Merged
pikann merged 1 commit into
masterfrom
fix/fix-error-cannot-upgrade
Oct 5, 2026
Merged

pikann merged 1 commit into
masterfrom
fix/fix-error-cannot-upgrade

Conversation

@pikann

@pikann pikann commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

upgrade.sh could exit with status 1 and no error message right after the "Proceed with upgrade?" prompt (reported in #546). get_env_var was grep … | head -1 | cut …; under set -euo pipefail, a grep miss makes the whole pipeline fail, so any bare X="$(get_env_var .env SOME_VAR)" assignment silently kills the script when SOME_VAR isn't in .env.

The first such read to hit a missing variable is AGENT_SERVER_IMAGE_AT_START (added with the image-pruning change in 2a89138), which runs immediately after the prompt. Other bare reads (STORAGE_PROVIDER, PACA_AI_AGENT_IMAGE, AGENT_SERVER_IMAGE) had the same exposure.

Changes

  • scripts/upgrade.sh, scripts/install.sh: get_env_var now prints nothing and succeeds when the variable is absent. install.sh had the same bug: a kept .env without AGENT_SERVER_IMAGE would die at the pre-pull step, before its existing fallback default could apply.
  • .github/workflows/scripts-pr-ci.yml: new upgrade-smoke step that upgrades a v0.18.0-style .env with no AGENT_SERVER_IMAGE, with Agent Runner both enabled and disabled (install.sh lets users opt out). It asserts the upgrade succeeds, the stack is brought up, and no image is invented or pulled.

Verification

  • Ran the new CI step locally against the workflow's fake docker: on the unfixed script it exits silently after "Target version" (same symptom as the issue); with the fix both variants pass.
  • bash -n passes on both scripts; the workflow YAML parses.
  • Not run locally: shellcheck (not installed) and the rest of the upgrade-smoke job; CI covers these.

Note on the root cause

The mechanism is reproduced on a synthetic .env, but I haven't seen the reporter's actual .env. A fresh v0.18.0 install.sh does write AGENT_SERVER_IMAGE, so the reporter's file presumably lacks it for another reason (hand-edited, manually created, kept from an older setup). Asking them for grep -c '^AGENT_SERVER_IMAGE=' .env and bash -x upgrade.sh output would confirm. The change is safe regardless, since it only affects variables that are allowed to be absent.

Refs #546

@pikann pikann changed the title fix(upgrade): add regression check for missing AGENT_SERVER_IMAGE in … fix(upgrade): don't exit silently when .env lacks AGENT_SERVER_IMAGE Oct 5, 2026

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes

  • get_env_var no longer fails on a missing variable — both scripts/upgrade.sh and scripts/install.sh now wrap the grep in { … || true; } before the head | cut pipeline, so an absent VAR yields an empty string instead of killing set -euo pipefail scripts at any X="$(get_env_var …)" assignment. The grouping also swallows the SIGPIPE grep gets when head -1 closes the pipe on duplicate keys.
  • Regression smoke test — .github/workflows/scripts-pr-ci.yml's upgrade-smoke job seeds a .env without AGENT_SERVER_IMAGE, runs the real upgrade.sh for PACA_AGENT_RUNNER=yes and no, and asserts the upgrade succeeds, compose … up runs, the runner flag is preserved, and no bare docker pull occurs.
  • Verified meaningful — both variants hit the missing-var read at scripts/upgrade.sh:383 before any Agent Runner branching, so the step fails against the pre-fix script.

I re-ran shellcheck --shell=bash scripts/install.sh scripts/upgrade.sh (a required lint job) and bash -n; both clean, and the workflow YAML parses. I also reproduced the original silent exit and confirmed the fix returns empty and proceeds. Not backfilling AGENT_SERVER_IMAGE is safe: deploy/docker-compose.prod.yml supplies ${AGENT_SERVER_IMAGE:-ghcr.io/paca-ai/paca-agent-server-goose:latest}, so the container still gets a default; only the best-effort pre-pull is skipped and the sandbox pulls lazily on first conversation.

Pullfrog  | View workflow run | Using deepseek-v4.1-flash (free via Pullfrog for OSS) | 𝕏

@pikann pikann mentioned this pull request Oct 5, 2026
@pikann
pikann merged commit 2ed7fe5 into master Oct 5, 2026
5 checks passed
@pikann
pikann deleted the fix/fix-error-cannot-upgrade branch October 5, 2026 16:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant