Skip to content

fix: keep followed logs alive across rollouts, restarts, and dropped streams - #785

Merged
nklmilojevic merged 8 commits into
mainfrom
fix/live-log-follow
Oct 6, 2026
Merged

nklmilojevic merged 8 commits into
mainfrom
fix/live-log-follow

Conversation

@nklmilojevic

Copy link
Copy Markdown
Owner

Summary

A followed kubelet log stream stopped for good once it ended. A container restart, an API server timeout, a dropped connection or waking the laptop left the view silent, with nothing to say it had stopped. Workload and service logs listed their pods once, so after a rollout the view went quiet and the new pods never appeared. Log streams are the most common complaint in k9s issues (k9s#1399, #1228 and #901). People still run stern next to their TUI because of this.

  • Reconnect from the last line. Followed streams reconnect using sinceTime from the last line shown. The API only accepts whole seconds there, so the server repeats up to a second of lines, and those already shown are skipped by timestamp. No line appears twice and none goes missing.
  • Read the pod when a stream ends. A [sofka] line marks a container restart or a recreated pod. The stream stops, with a reason, when the pod is deleted or finished, or when a container exits under restartPolicy: Never or OnFailure. A 4xx refusal is reported once and not retried. Reconnects back off from 1 s up to 15 s while a stream returns nothing new.
  • Follow rollouts. Workload and service logs watch the label selector instead of listing pods once. A new pod gets a following new pod line and is shown from its first line. Pods that already existed keep the configured tail.
  • Survive sleep. The existing wake detector now also reconnects log streams and the pod watch. Opening a stream times out after 30 s and the pod check after 10 s, so a connection left dead by sleep can't hang the view.

Test plan

  • just check: fmt, clippy -D warnings, 1949 tests
  • New tests in src/app/tests/log_follow.rs, all starting from l: a resume after restart with no repeated lines, a deleted pod, a 403 reported once, a rollout adding a pod, and waking from sleep
  • Unit tests for the skip-repeats logic and the restart, recreate and finish checks
  • Manual: l on a Deployment, then kubectl rollout restart; new pods join the view
  • Manual: follow a pod and kill its container; a restart line appears and logs continue
  • Manual: follow logs, sleep the laptop, wake it; the stream resumes

…streams

A followed kubelet log stream stopped for good when it ended: a container
restart, an API server timeout, a dropped connection, or waking the laptop
left the view silent with no indication. Workload and service logs listed
their pods once, so a rollout replaced every streamed pod and the view went
quiet while the new pods were never shown.

Followed streams now reconnect from the last delivered line. sinceTime only
has second resolution, so the replayed second is de-duplicated by timestamp.
When a stream ends, the pod is read to report a restart or recreation, or to
stop once the pod is deleted or finished. Refusals (4xx) are reported once
instead of retried. Selector logs watch the pods instead of listing them, so
pods a rollout creates are followed from their first line. Waking from sleep
reconnects log streams as it already restarts the table watch.
@kritikal-github

kritikal-github Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Re-runKritika Review

Makes rollout log-follow test ordering deterministic with future timestamps

No findings

Confidence 5/5 · medium risk: The risk rests on changes to application log-follow behavior and Kubernetes watch handling; the review reported no findings, and I found no additional issue.

Incremental review of the changes since b581b30.

Summary

The update gives both pod log lines future timestamps, so a rollout notice retains a stable position even if it arrives before the first line. The test continues to verify that the new pod is followed without an initial tail or lookback.

What's good

  • Accounts for notices arriving before either pod's log lines.

1 diff file(s) left out of the prompt to fit its budget: src/app/tests/log_follow.rs.
41 context chunk(s) left out of the prompt to fit its budget.

Reviewed 56d5009 by kritika with chatgpt/gpt-6-sol.

Comment thread src/app/log_follow.rs Outdated
Comment thread src/app/log_follow.rs Outdated
Comment thread src/app/log_follow.rs
Comment thread src/app/log_follow.rs
@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Rewrites log streaming to handle reconnects and pod lifecycle.

The PR appears safe to merge based on this review; the latest change introduces no actionable issue.

Summary

The PR reconnects followed pod logs, checks pod identity and status around reconnects, and watches workload and service selectors so newly created pods join the view.

  • The only change since the previous review moves two rollout-test log timestamps into the future to make notice ordering deterministic.
  • No new actionable issue was identified.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Open followed logs] --> B[Stream pod logs]
  B --> C{Stream ends or machine wakes}
  C --> D[Check pod identity and status]
  D -->|Same active pod| E[Resume from last line]
  D -->|Recreated pod| F[Read from first line]
  D -->|Gone or finished| G[End with notice]
  E --> B
  F --> B
  H[Workload or service selector watch] -->|New matching pod| A
Loading

Reviews (8) · Last reviewed commit: "test: date the rollout log lines in the ..."

Comment thread src/app/log_follow.rs Outdated
Comment thread src/app/log_follow.rs
Comment thread src/app/log_follow.rs Outdated
Comment thread src/app/logs.rs Outdated
Comment thread src/app/log_follow.rs Outdated
Review found gaps in the log follow lifecycle. A pod that stopped matching
the selector kept streaming into the view because a watch deletion only
forgot its key. A recreated pod inherited the old instance's resume
position, so its early lines could be skipped. A 404 when reopening a stream
was treated as a refusal and skipped the pod check that reports deletion. A
refused pod watch was retried forever after the first sync. A log request
stuck on a connection that died during sleep ignored the wake signal until
its 30-second timeout. Single-pod streams also learned the pod's UID only
after their first end, so an early recreation went unnoticed.
Comment thread src/app/logs.rs Outdated
Comment thread src/app/log_follow.rs Outdated
…replacement

Log requests go by pod name, so a marked pod replaced by one with the same
name kept streaming the replacement, even though the marked set is fixed
when the view opens. Marked streams are now tied to the pod's UID and end
with a notice when it is replaced.

When a pod was replaced before the first log request answered, that request
already read the replacement. The later UID check then reset the stream and
read the replacement from its start again, so its lines showed twice. The
reset now happens only when the lines shown so far predate the new pod.
Comment thread src/app/log_follow.rs
Comment thread src/app/log_follow.rs Outdated
Log requests go by pod name. A marked or selector pod replaced before its
stream opened was streamed under the old pod's identity, and an open stream
never reached the end-of-stream check that would notice. The replacement
check after a stream also compared log timestamps, written on the node, with
the pod's creation time, set by the API server, so clock skew could drop or
repeat the replacement's lines.

When the pod's UID is known, its identity is now checked before every log
request. A replacement is found before any of its lines are shown, so a
recreated pod is always read from its start, and a marked or selector pod
ends instead of streaming the replacement.
Comment thread src/app/log_follow.rs Outdated
Comment thread src/app/log_follow.rs Outdated
The pod check added before each log request waited out its 10-second
timeout on a connection that died during sleep, delaying the return of live
logs after a wake. The stream's start time was also taken after that check,
so a container that restarted while it was pending looked like it restarted
before the stream and lost its restart notice.

Both pod checks now give up on a wake, and the start time is taken before
the check.
Comment thread src/app/log_follow.rs
Comment thread src/app/log_follow.rs
Comment thread src/app/log_follow.rs Outdated
Comment thread src/app/log_follow.rs Outdated
A failed or timed-out pod check let a marked or selector stream open its
log request anyway, which could attach to a replacement pod with the same
name. Pinned streams now retry the check instead. A refused check still
opens the stream unchecked, since the check could never pass.

A container waiting to start was retried every second without reading the
pod, so a pod that failed or went away in the meantime kept the stream
retrying forever. The wait now reads the pod and backs off up to five
seconds.

Restart notices could be lost when a wake interrupted the pod read after a
stream, or misattributed by comparing whole seconds. The check before each
request now records the restart count, so restarts are found by count, and
the read after a stream is retried after a wake.
Comment thread src/app/log_follow.rs
Where RBAC allowed reading logs but not getting pods, the identity check was
refused and a marked or selector stream opened by name anyway, so a
replacement with the same name could be shown in its place. The check now
falls back to listing the pod by name, which the table and the selector watch
already need. If both are refused, a pinned stream ends with a notice instead
of opening unchecked.
A notice that arrives before any log line sorts at the current time. The
rollout test dated its lines today, so after 10:00 UTC the new-pod notice
sorted after the new pod's line and CI failed depending on when it ran.
@nklmilojevic
nklmilojevic merged commit d7b896b into main Oct 6, 2026
5 checks passed
@nklmilojevic
nklmilojevic deleted the fix/live-log-follow branch October 6, 2026 10:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant