Skip to content

Clarify mining logs: real job-stop reasons and per-type worker numbering - #108

Merged
illuzen merged 1 commit into
mainfrom
illuzen/better-logs
Sep 18, 2026
Merged

illuzen merged 1 commit into
mainfrom
illuzen/better-logs

Conversation

@illuzen

@illuzen illuzen commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Overview

Miners (especially newcomers) were confused by "job cancelled" logs — which almost always just mean a new block arrived — and by inconsistent numbering where the same GPU appeared as "GPU worker 5" in service logs but "GPU 0" in engine logs.

What changed

miner-service

  • WorkerPool now records why a job stopped via a new JobStopReason enum (NewBlock, SolutionFound, ConnectionLost), so worker completion lines state the actual reason instead of assuming "new block" for every stop:
    • CPU worker 2 stopped (new block received): 2170000 hashes in 4.90s (442.83 KH/s)
    • GPU worker 0 stopped (solution already found): ...
    • CPU worker 1 stopped (node connection lost): ...
    • CPU worker 3 finished (nonce range exhausted): ...
  • Renamed WorkerPool::cancel() → stop_current_job(reason).
  • Workers are numbered per engine type (CPU 0..N, GPU 0..M) instead of one global counter, so GPU worker numbers match how users think about their hardware; WorkerResult.thread_id → worker_id.
  • Job lifecycle logs reworded: new job arrival is now 🧱 New block to mine: job 1393 (header 0x...); solution submission is ✅ Job 1393 solved: ... - submitting solution to node. Renamed result_sent_for_current_job → solution_submitted_for_current_job.

engine-gpu / engine-cuda

  • Engine logs now say GPU device N / CUDA device N (instead of bare GPU N) to distinguish hardware from worker threads.
  • The wgpu engine's per-cancellation info log (GPU 0 cancelled before batch 5) is reworded to GPU device 0 search stopped after 5 batches and demoted to debug — the service now emits the clear info-level line with the actual stop reason.

No mining behavior changes: log text, log levels, and internal renames only. Engine-level EngineStatus::Cancelled / CancelCheck APIs are intentionally unchanged (accurate mechanical terms shared across engine crates and benchmarks).

Validation

  • ./clippy.sh (taplo fmt + cargo fmt + clippy -D warnings): clean
  • cargo test --workspace --locked: 67 tests pass, 0 failures
  • Manually traced the log flow for the new-block, solution-found, and connection-lost paths

Risks and mitigations

  • JobStopReason is read by workers after they observe a job-ID change; a tiny race could log a stale reason if two stops happen back-to-back. Log-text-only impact, no functional risk.
  • With multiple GPUs, per-type worker numbers may not align one-to-one with device indices (round-robin assignment on first search); the existing Worker thread assigned to GPU device N info line shows the mapping.

Follow-ups

None.

Made with Cursor

Miners were confused by "cancelled" logs when a new block arrived, and
by mismatched numbering between service workers and engine devices
("GPU worker 5" vs "GPU 0").

- Track why a job stopped (JobStopReason: new block, solution found,
  connection lost) in WorkerPool and log it in worker completion lines,
  e.g. "GPU worker 0 stopped (new block received): ..."
- Rename WorkerPool::cancel() to stop_current_job(reason)
- Number workers per engine type (CPU 0..N, GPU 0..M) instead of one
  global counter; rename WorkerResult.thread_id to worker_id
- Engine logs now say "GPU device N" / "CUDA device N" to distinguish
  hardware from worker threads
- Reword job lifecycle logs: new job arrival is "🧱 New block to mine",
  solution submission is "✅ Job N solved ... submitting solution"
- Demote wgpu engine's per-cancellation info log to debug; the service
  now emits the clear info-level line with the actual reason

No mining behavior changes; log text, log levels, and internal renames
only.

Co-authored-by: Cursor <cursoragent@cursor.com>

@n13 n13 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer model: GPT 5.6 Sol

APPROVE — no blocking findings.

The stop reason is stored before each job-generation change, every stop site supplies the appropriate reason, and the per-engine worker numbering is carried consistently through worker results and stale-result logs. The engine changes are limited to log wording and log level.

Validation:

  • taplo format --check --config taplo.toml — passed
  • cargo +nightly fmt --all -- --check — passed
  • cargo test --workspace --locked — passed (67 tests)
  • cargo clippy --workspace --all-targets --all-features --locked -- -D warnings — passed
  • GitHub CI — all six checks passed at a2007f4928ef48935f2d7bfdb51f48e1d3393af4

The PR's documented rapid-successive-stop race can still make the displayed stop reason stale, but it does not affect cancellation or mining behavior and is not blocking for this logging-focused change.

@n13 n13 removed the bot-review label Sep 17, 2026
@n13

n13 commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

LGTM

@illuzen
illuzen merged commit c7838cb into main Sep 18, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants