Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
218 changes: 218 additions & 0 deletions .claude/skills/heartbeat/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,218 @@
---
name: heartbeat
description: Work with the gProfiler heartbeat system for dynamic profiling control. Use when the user asks about heartbeat mode, Performance Studio integration, or command-driven profiling.
---

## gProfiler Heartbeat System

The heartbeat system enables centralized profiling control where Performance Studio can dynamically issue start/stop commands to gProfiler agents.

### System Architecture

```
┌─────────────────────┐ Heartbeat ┌──────────────────────┐
│ Performance Studio │ ◄──────────────► │ gProfiler Agent │
│ Backend │ Commands │ │
└─────────────────────┘ ────────────────► └──────────────────────┘
```

### Running in Heartbeat Mode

**Basic:**
```bash
python gprofiler/main.py \
--enable-heartbeat-server \
--upload-results \
--token "your-token" \
--service-name "web-service" \
--api-server "http://performance-studio:8000" \
--heartbeat-interval 30 \
--output-dir /tmp/profiles \
--verbose
```

**Production:**
```bash
export GPROFILER_TOKEN="my_token"
export GPROFILER_SERVICE="your-service-name"
export GPROFILER_SERVER="http://localhost:8080"

/opt/gprofiler/gprofiler \
--enable-heartbeat-server \
-u \
--token=$GPROFILER_TOKEN \
--service-name=$GPROFILER_SERVICE \
--api-server $GPROFILER_SERVER \
--dont-send-logs \
--server-upload-timeout 10 \
-c \
--disable-metrics-collection \
--java-safemode= \
--heartbeat-interval 30 \
-d 60 \
--java-no-version-check
```

`--server-host` still exists as a deprecated alias, but prefer `--api-server`.

### Required flags

Current `main.py` validation requires heartbeat mode to include:

- `--enable-heartbeat-server`
- `--upload-results`
- `--token`
- `--service-name`

Use the skill to explain or debug this mode only in terms of the current flags above.

### Command Flow

```
1. User submits profiling request to backend
↓
2. Backend creates command with unique ID
↓
3. Agent sends heartbeat to backend
↓
4. Backend responds with pending command
↓
5. Agent checks idempotency (skip if already received)
↓
6. Agent enqueues command in priority queue
↓
7. Agent executes command (start/stop profiling)
↓
8. Agent reports completion to backend
```

### Command Priority Queues

| Queue | Purpose | Max Size |
|-------|---------|----------|
| `stop_queue` | Immediate stop commands | 1 |
| `adhoc_queue` | Single-run start commands | 10 |
| `continuous_queue` | Long-running start commands | 1 |

Priority: `stop > adhoc > continuous`

The current implementation lives under `gprofiler/dynamic_profiling_management/`. Do not refer users to `gprofiler/command_control.py`; that path is stale.

### API Endpoints

**Submit Profiling Request:**
```bash
curl -X POST http://localhost:8000/api/metrics/profile_request \
-H "Content-Type: application/json" \
-d '{
"service_name": "web-service",
"command_type": "start",
"duration": 60,
"frequency": 11,
"profiling_mode": "cpu",
"target_hostnames": ["host1", "host2"]
}'
```

**Stop Profiling:**
```bash
curl -X POST http://localhost:8000/api/metrics/profile_request \
-H "Content-Type: application/json" \
-d '{
"service_name": "web-service",
"command_type": "stop",
"stop_level": "host",
"target_hostnames": ["host1"]
}'
```

### PerfSpect Hardware Metrics

Enable Intel PerfSpect for hardware metrics:
```bash
curl -X POST http://localhost:8000/api/metrics/profile_request \
-H "Content-Type: application/json" \
-d '{
"service_name": "web-service",
"command_type": "start",
"duration": 60,
"additional_args": {
"enable_perfspect": true
}
}'
```

Requirements:
- Linux x86_64 (Intel architecture)
- Root access
- Internet for auto-install

### Key Files

```
gprofiler/main.py # CLI + heartbeat flag validation
gprofiler/dynamic_profiling_management/heartbeat.py # Polling and command handling
gprofiler/dynamic_profiling_management/command_control.py # Queue logic and priority
gprofiler/dynamic_profiling_management/continuous.py # Continuous slot
gprofiler/dynamic_profiling_management/ad_hoc.py # Ad-hoc slot
tests/test_heartbeat_system.py # Heartbeat flow validation
docs/HEARTBEAT_SYSTEM_README.md # Full documentation
```

### Testing heartbeat changes

Use the smallest useful validation first:

```bash
# Focused heartbeat test
sudo python3 -m pytest -v tests/test_heartbeat_system.py

# Lightweight broader regression
sudo ./tests/test.sh --executable
```

For local end-to-end testing against a backend, the repo docs describe this sequence:

1. Start the Performance Studio backend.
2. Run `python tests/run_heartbeat_agent.py`
3. Submit commands with `python tests/test_heartbeat_system.py --live`

Prefer the existing docs/test scripts over inventing custom heartbeat harnesses.

### Troubleshooting

**Agent not receiving commands:**
- Check network connectivity
- Verify authentication token
- Check service name matching

**Commands not executing:**
- Check agent logs for errors
- Verify command parameters
- Check system permissions

**PerfSpect not working:**
- Verify Linux x86_64 platform
- Check root permissions
- Check `/tmp/gprofiler_perfspect/perfspect/`

### CLI Options Reference

```bash
--enable-heartbeat-server # Enable heartbeat mode
--heartbeat-interval 30 # Heartbeat frequency (seconds)
--api-server URL # Backend server URL
--server-host URL # Deprecated alias for --api-server
--upload-results # Required for heartbeat mode
--token TOKEN # Authentication token
--service-name NAME # Service identifier
--enable-hw-metrics-collection # Enable PerfSpect
--perfspect-path PATH # PerfSpect binary path
```

### Review points for heartbeat work

- Preserve queue semantics: `stop > adhoc > continuous`
- Preserve idempotency; do not allow the same command to execute twice
- Avoid moving heartbeat logic into `main.py` if `dynamic_profiling_management/` is sufficient
- Add targeted heartbeat tests before broader regression runs
142 changes: 142 additions & 0 deletions METRICS_README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,142 @@
# gProfiler Metrics Implementation

## Overview

gProfiler metrics system provides comprehensive error monitoring and observability by sending structured metrics to Pinterest's MetricAgent.

## Architecture

### Core Components

```
MetricsHandler (Singleton)
├── decorate_metric_name() # Hierarchical naming
├── build_enriched_tags() # System + user tags
├── format_metric_message() # Goku protocol formatting
└── send_metric() # TCP transmission
```

### Design Principles

- **Single Responsibility**: Each method has one clear purpose
- **Resource Efficiency**: Singleton pattern ensures one TCP connection
- **Clean API**: Simple, intuitive method names
- **Testability**: Pure functions with clear inputs/outputs
- **Robustness**: Never crashes main application

## Error Types & Purpose

### 🔍 **Error Segregation Strategy**

| Error Type | Purpose | When Used |
|------------|---------|-----------|
| `process_profiler_failure` | **Individual process profiler crashes** | Java/Python/Native profiler dies |
| `perf_failure` | **System perf command failures** | `perf record` command fails |
| `profiling_run_failure` | **Entire profiling cycle crashes** | Complete profiling session fails |
| `upload_error` | **Profile upload errors** | Network/server upload errors |
| `api_error` | **HTTP 4xx/5xx from API server** | Server returns error status |
| `request_exception` | **Network connection failures** | TCP/DNS/connection errors |

**Why segregate?** Different error types require different:
- **Operational responses** (restart profiler vs fix network)
- **Alert routing** (infra team vs app team)
- **SLA tracking** (profiler reliability vs upload reliability)

## Usage

### Basic Usage

```python
from gprofiler.metrics_publisher import (
MetricsHandler,
ERROR_TYPE_PROCESS_PROFILER_FAILURE,
COMPONENT_SYSTEM_PROFILER,
SEVERITY_ERROR,
get_current_method_name
)

# Singleton - same instance everywhere
handler = MetricsHandler('tcp://localhost:18126', 'gprofiler')

# Send error metric
handler.send_error_metric(
error_type=ERROR_TYPE_PROCESS_PROFILER_FAILURE,
error_message="Java profiler crashed during heap analysis",
category=COMPONENT_SYSTEM_PROFILER,
severity=SEVERITY_ERROR,
extra_tags={
'method_name': get_current_method_name(),
'profiler_name': 'java',
'failure_reason': 'out_of_memory'
}
)
```

### Configuration

```python
# CLI Arguments
parser.add_argument('--enable-publish-metrics', action='store_true')
parser.add_argument('--metrics-server-url', default='tcp://localhost:18126')
parser.add_argument('--service-name', default='gprofiler')

# Initialization
if args.enable_publish_metrics:
handler = MetricsHandler(args.metrics_server_url, args.service_name)
else:
handler = NoopMetricsHandler() # Safe no-op when disabled
```

## Metric Format

### Hierarchical Naming

```
gprofiler.{category}.{error_type}.error
```

**Examples:**
- `gprofiler.system_profiler.process_profiler_failure.error`
- `gprofiler.api_client.upload_error.error`
- `gprofiler.gprofiler_main.profiling_run_failure.error`

### Tags (Metadata)

**System Tags (automatic):**
```json
{
"service": "gprofiler",
"hostname": "prod-server-01",
"component": "system_profiler",
"severity": "error",
"os_type": "linux",
"python_version": "3.8"
}
```

**Runtime Tags (when gProfiler is active):**
```json
{
"run_id": "gprofiler-1761058503",
"cycle_id": "42"
}
```
*Note: run_id and cycle_id are only present when gProfiler is actively profiling*

**User Tags (custom):**
```json
{
"method_name": "GProfiler._snapshot",
"profiler_name": "java",
"failure_reason": "timeout",
"duration_ms": "5000"
}
```

### Goku Protocol Message

```
put gprofiler.system_profiler.process_profiler_failure.error 1761094217 1 service=gprofiler hostname=hostname component=system_profiler severity=error os_type=linux python_version=3.8 method_name=GProfiler._snapshot profiler_name=java
```

*Note: Actual message includes all system tags (service, hostname, component, severity, os_type, python_version) plus any user tags. Runtime tags (run_id, cycle_id) are added when gProfiler is actively profiling.*
Loading
Loading