Repository navigation
fix: keep followed logs alive across rollouts, restarts, and dropped streams - #785
Conversation
…streams A followed kubelet log stream stopped for good when it ended: a container restart, an API server timeout, a dropped connection, or waking the laptop left the view silent with no indication. Workload and service logs listed their pods once, so a rollout replaced every streamed pod and the view went quiet while the new pods were never shown. Followed streams now reconnect from the last delivered line. sinceTime only has second resolution, so the replayed second is de-duplicated by timestamp. When a stream ends, the pod is read to report a restart or recreation, or to stop once the pod is deleted or finished. Refusals (4xx) are reported once instead of retried. Selector logs watch the pods instead of listing them, so pods a rollout creates are followed from their first line. Waking from sleep reconnects log streams as it already restarts the table watch.
|
|
Review found gaps in the log follow lifecycle. A pod that stopped matching the selector kept streaming into the view because a watch deletion only forgot its key. A recreated pod inherited the old instance's resume position, so its early lines could be skipped. A 404 when reopening a stream was treated as a refusal and skipped the pod check that reports deletion. A refused pod watch was retried forever after the first sync. A log request stuck on a connection that died during sleep ignored the wake signal until its 30-second timeout. Single-pod streams also learned the pod's UID only after their first end, so an early recreation went unnoticed.
…replacement Log requests go by pod name, so a marked pod replaced by one with the same name kept streaming the replacement, even though the marked set is fixed when the view opens. Marked streams are now tied to the pod's UID and end with a notice when it is replaced. When a pod was replaced before the first log request answered, that request already read the replacement. The later UID check then reset the stream and read the replacement from its start again, so its lines showed twice. The reset now happens only when the lines shown so far predate the new pod.
Log requests go by pod name. A marked or selector pod replaced before its stream opened was streamed under the old pod's identity, and an open stream never reached the end-of-stream check that would notice. The replacement check after a stream also compared log timestamps, written on the node, with the pod's creation time, set by the API server, so clock skew could drop or repeat the replacement's lines. When the pod's UID is known, its identity is now checked before every log request. A replacement is found before any of its lines are shown, so a recreated pod is always read from its start, and a marked or selector pod ends instead of streaming the replacement.
The pod check added before each log request waited out its 10-second timeout on a connection that died during sleep, delaying the return of live logs after a wake. The stream's start time was also taken after that check, so a container that restarted while it was pending looked like it restarted before the stream and lost its restart notice. Both pod checks now give up on a wake, and the start time is taken before the check.
A failed or timed-out pod check let a marked or selector stream open its log request anyway, which could attach to a replacement pod with the same name. Pinned streams now retry the check instead. A refused check still opens the stream unchecked, since the check could never pass. A container waiting to start was retried every second without reading the pod, so a pod that failed or went away in the meantime kept the stream retrying forever. The wait now reads the pod and backs off up to five seconds. Restart notices could be lost when a wake interrupted the pod read after a stream, or misattributed by comparing whole seconds. The check before each request now records the restart count, so restarts are found by count, and the read after a stream is retried after a wake.
Where RBAC allowed reading logs but not getting pods, the identity check was refused and a marked or selector stream opened by name anyway, so a replacement with the same name could be shown in its place. The check now falls back to listing the pod by name, which the table and the selector watch already need. If both are refused, a pinned stream ends with a notice instead of opening unchecked.
A notice that arrives before any log line sorts at the current time. The rollout test dated its lines today, so after 10:00 UTC the new-pod notice sorted after the new pod's line and CI failed depending on when it ran.
Summary
A followed kubelet log stream stopped for good once it ended. A container restart, an API server timeout, a dropped connection or waking the laptop left the view silent, with nothing to say it had stopped. Workload and service logs listed their pods once, so after a rollout the view went quiet and the new pods never appeared. Log streams are the most common complaint in k9s issues (k9s#1399, #1228 and #901). People still run stern next to their TUI because of this.
sinceTimefrom the last line shown. The API only accepts whole seconds there, so the server repeats up to a second of lines, and those already shown are skipped by timestamp. No line appears twice and none goes missing.[sofka]line marks a container restart or a recreated pod. The stream stops, with a reason, when the pod is deleted or finished, or when a container exits underrestartPolicy: NeverorOnFailure. A 4xx refusal is reported once and not retried. Reconnects back off from 1 s up to 15 s while a stream returns nothing new.following new podline and is shown from its first line. Pods that already existed keep the configured tail.Test plan
just check: fmt, clippy-D warnings, 1949 testssrc/app/tests/log_follow.rs, all starting froml: a resume after restart with no repeated lines, a deleted pod, a 403 reported once, a rollout adding a pod, and waking from sleeplon a Deployment, thenkubectl rollout restart; new pods join the view