Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/copilot-instructions.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ Multi-package colcon workspace under `src/`:
| Package | Purpose |
|---------|---------|
| `ros2_medkit_gateway` | HTTP gateway - REST server, discovery, entity management, handlers, plugin framework |
| `ros2_medkit_fault_manager` | Fault aggregation with SQLite, AUTOSAR DEM-style debounce, rosbag/snapshot capture |
| `ros2_medkit_fault_manager` | Fault records per (code, reporting source) in SQLite, AUTOSAR DEM-style debounce, rosbag/snapshot capture |
| `ros2_medkit_fault_reporter` | Client library for nodes to report faults |
| `ros2_medkit_diagnostic_bridge` | Bridges `/diagnostics` topic to fault manager |
| `ros2_medkit_serialization` | Runtime JSON <-> ROS 2 message serialization via dynmsg |
Expand Down
42 changes: 31 additions & 11 deletions docs/api/messages.rst
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,12 @@ Messages
Fault.msg
~~~~~~~~~

Core fault data model representing an aggregated fault condition.
Core fault data model representing one fault record.

A record is identified by the pair (``fault_code``, owning reporting source). The owner is
the ``source_id`` a ``ReportFault`` call carried. Two sources reporting one ``fault_code``
are two records, each with its own status, debounce counter, ``occurrence_count``, severity
and timestamps, and each cleared on its own.

.. code-block:: text

Expand Down Expand Up @@ -49,7 +54,8 @@ Core fault data model representing an aggregated fault condition.
# Current fault status (PREFAILED, PREPASSED, CONFIRMED, HEALED, CLEARED)
string status

# List of source identifiers that have reported this fault
# The reporting source that owns this record, as a one-element list. Together
# with fault_code it identifies the record.
string[] reporting_sources

**Severity Constants:**
Expand Down Expand Up @@ -146,14 +152,18 @@ prefix (for example ``/robot1/fault_manager/events``).
MutedFaultInfo.msg
~~~~~~~~~~~~~~~~~~

Information about correlated (muted) symptom faults.
Information about a correlated (muted) symptom record. One entry per muted RECORD:
a root cause mutes the symptoms of its own reporting source only, so two sources
muted on one ``fault_code`` produce two entries carrying that code, told apart by
``source_id``.

.. code-block:: text

string fault_code # The muted symptom's fault code
string root_cause_code # Root cause that triggered muting
string rule_id # Correlation rule ID that matched
uint32 delay_ms # Time delay from root cause [ms]
string source_id # Reporting source that owns the muted record

ClusterInfo.msg
~~~~~~~~~~~~~~~
Expand Down Expand Up @@ -189,7 +199,8 @@ Report a fault event to the FaultManager.
uint8 event_type # EVENT_FAILED (0) or EVENT_PASSED (1)
uint8 severity # Fault.SEVERITY_* constant (for FAILED events)
string description # Human-readable description
string source_id # Fully qualified node name (e.g., "/powertrain/temp_sensor")
string source_id # Fully qualified node name (e.g., "/powertrain/temp_sensor").
# Owns the record: (fault_code, source_id) is the record identity.

**Response:**

Expand All @@ -215,23 +226,31 @@ Report a fault event to the FaultManager.
ClearFault.srv
~~~~~~~~~~~~~~

Clear/acknowledge a fault.
Clear/acknowledge one fault record.

**Request:**

.. code-block:: text

string fault_code # Fault code to clear
bool skip_correlation_auto_clear # Opt out of correlation cascade clear
string source_id # Reporting source that owns the record

**Response:**

.. code-block:: text

bool success # True if fault was found and cleared
bool success # True if the record was found and cleared
string message # Status message or error description
string[] auto_cleared_codes # Symptoms auto-cleared with root cause

``source_id`` names the owner of the record to clear. Leaving it empty is unscoped: the
call applies only when exactly one record carries the ``fault_code``, and fails otherwise
with ``success=false`` and a ``message`` beginning ``ambiguous:`` that lists the owners.
An ambiguous call clears nothing. ``GetFault``, ``GetSnapshots`` and ``GetRosbag`` carry
the same field and resolve the same way. On ``GetRosbag`` it scopes the ``fault_code``
lookup only, and the ``recording_id`` path ignores it.

When ``skip_correlation_auto_clear`` is ``false`` (default), clearing a
root-cause fault also clears every symptom that the correlation engine
attributes to it via ``auto_clear_with_root`` rules; the cleared
Expand All @@ -246,11 +265,12 @@ cluster-wide clearing still works.

.. note::

Added in ``ros2_medkit_msgs`` post-0.4.0. Adding a request field
changes the service type hash, so out-of-tree callers that invoke
``/fault_manager/clear_fault`` directly via ``ros2 service call`` or
a generated client must rebuild against the new ``ros2_medkit_msgs``
release to keep talking to ``fault_manager``.
``skip_correlation_auto_clear`` was added in ``ros2_medkit_msgs`` post-0.4.0,
and ``source_id`` after it (the same field was added to ``GetFault``,
``GetSnapshots`` and ``GetRosbag``). Adding a request field changes the
service type hash, so out-of-tree callers that invoke those services directly
via ``ros2 service call`` or a generated client must rebuild against the new
``ros2_medkit_msgs`` release to keep talking to ``fault_manager``.

ListFaults.srv
~~~~~~~~~~~~~~
Expand Down
46 changes: 25 additions & 21 deletions docs/config/fault-manager.rst
Original file line number Diff line number Diff line change
@@ -1,7 +1,8 @@
Fault Manager Configuration
===========================

The ``ros2_medkit_fault_manager`` node aggregates and manages faults from multiple sources.
The ``ros2_medkit_fault_manager`` node keeps and manages one fault record per
(``fault_code``, reporting source) pair.
This page documents all configuration parameters.

.. contents:: Table of Contents
Expand Down Expand Up @@ -114,25 +115,26 @@ A **near miss** is a FAILED report that moved the debounce counter without the f
CONFIRMED - the fault nearly happened. PASSED reports move the counter in the healing direction
(the fault receding) and are not near misses.

The fault manager appends one entry per near miss to a per-fault-code series, holding the
timestamp, the counter value after the report, the confirmation threshold, the severity, the
reporting source and the fault status the report left behind.
The fault manager appends one entry per near miss to the series of the fault record it moved (one
series per fault code and reporting source), holding the timestamp, the counter value after the
report, the confirmation threshold, the severity, the reporting source and the fault status the
report left behind.

The status matters when reading the series. The HEALED latch holds the status all the way from the
healing threshold down to the confirmation threshold, so reports on the way back into a fault that
does confirm are also near misses by the definition above. Entries recording ``PREFAILED`` are
approaches from a resting state; entries recording ``HEALED`` are a counter walking back down under
the latch. The recorded confirmation threshold belongs to the reporting source, while the counter
is shared by all sources of that fault code, so with per-entity thresholds it is not on its own the
distance to confirmation. The series is **retained when the fault is cleared**, because acknowledging one
fault cycle must not erase how often that code approached confirmation across cycles.
the latch. The recorded confirmation threshold and the counter both belong to the record's own
reporting source, so with per-entity thresholds each entry still reads as the record's distance to
confirmation. The series is **retained when the fault is cleared**, because acknowledging one
fault cycle must not erase how often that record approached confirmation across cycles.

.. code-block:: yaml

fault_manager:
ros__parameters:
near_miss:
max_per_fault: 200 # Entries kept per fault code (0 = unlimited)
max_per_fault: 200 # Entries kept per fault record (0 = unlimited)

.. list-table::
:header-rows: 1
Expand All @@ -143,7 +145,7 @@ fault cycle must not erase how often that code approached confirmation across cy
- Description
* - ``near_miss.max_per_fault``
- ``200``
- Near-miss entries retained per fault code. When the bound is reached the **oldest**
- Near-miss entries retained per fault record. When the bound is reached the **oldest**
entries are evicted, the same direction as ``snapshots.max_per_fault`` and the rosbag cap:
a series frozen at boot says nothing about whether the rate of near misses is changing.
Set to 0 for unlimited, accepting growth with the reporting rate.
Expand Down Expand Up @@ -208,8 +210,8 @@ threshold overrides:

- The ``source_id`` is the identifier passed in ``ReportFault`` service requests, typically the
fully qualified name of the reporting ROS 2 node (e.g., ``/sensors/lidar/front_node``).
You can inspect actual ``source_id`` values in the ``reporting_sources`` field of existing
faults via ``GET /api/v1/faults``.
It also owns the record it creates, so you can inspect actual ``source_id`` values in the
one-element ``reporting_sources`` field of existing faults via ``GET /api/v1/faults``.
- The ``source_id`` from ``ReportFault`` requests is matched against configured prefixes.
- The **longest matching prefix** wins. For example, ``/sensors/lidar/front`` matches
``/sensors/lidar`` over ``/sensors``.
Expand All @@ -219,9 +221,11 @@ threshold overrides:

.. note::

When multiple entities report the same ``fault_code``, each event applies the
thresholds resolved from that event's ``source_id``. This means the debounce
behavior follows the reporting entity, not the fault.
When multiple entities report the same ``fault_code``, each of them owns its own
record and each event applies the thresholds resolved from that event's
``source_id``. The debounce counter belongs to the same record, so an entity's
configured band governs exactly the counter its own reports move: one entity's
reports can neither confirm nor heal another entity's fault.

``auto_confirm_after_sec`` is global-only and cannot be overridden per-entity.
Critical faults skip debounce and confirm on their first occurrence; that is
Expand All @@ -246,8 +250,8 @@ Basic Snapshot Settings
max_message_size: 65536 # Max message size in bytes (64KB)
default_topics: [] # Topics to capture for all faults
config_file: "" # Path to YAML config file
recapture_cooldown_sec: 60.0 # Min seconds between snapshot captures per fault
max_per_fault: 10 # Max snapshots stored per fault code (0 = unlimited)
recapture_cooldown_sec: 60.0 # Min seconds between snapshot captures per fault record
max_per_fault: 10 # Max snapshots stored per fault record (0 = unlimited)
capture_pool_size: 2 # Max concurrent capture threads (>= 1)
capture_queue_depth: 16 # Max pending captures before policy applies (>= 1)
capture_queue_full_policy: reject_newest # reject_newest | drop_oldest
Expand Down Expand Up @@ -284,11 +288,11 @@ Basic Snapshot Settings
- Path to YAML file with fault-specific snapshot configurations.
* - ``snapshots.recapture_cooldown_sec``
- ``60.0``
- Minimum seconds between snapshot captures for the same fault code.
- Minimum seconds between snapshot captures for the same fault record.
Prevents snapshot storms when a fault is reported repeatedly. Set to 0 to disable.
* - ``snapshots.max_per_fault``
- ``10``
- Maximum number of snapshot rows stored per fault code. One confirmation
- Maximum number of snapshot rows stored per fault record. One confirmation
writes one row per configured topic, and those rows are evicted together:
past the limit the OLDEST capture set is dropped whole. A capture larger
than the cap is kept anyway rather than torn, since half a freeze frame is
Expand Down Expand Up @@ -339,7 +343,7 @@ Capture continuous rosbag recordings around fault events.
max_buffer_mb: 256 # Ring-buffer RAM cap
max_bag_size_mb: 50 # Max size per bag file
max_total_storage_mb: 500 # Max total storage
max_bags_per_fault: 1 # Recordings kept per fault code
max_bags_per_fault: 1 # Recordings kept per fault record
auto_cleanup: true # Auto-delete old bags

.. list-table::
Expand Down Expand Up @@ -416,7 +420,7 @@ Capture continuous rosbag recordings around fault events.
whole burst's bag at a time (oldest first).
* - ``rosbag.max_bags_per_fault``
- ``1``
- How many recordings one fault code keeps. Past the cap the oldest is
- How many recordings one fault record keeps. Past the cap the oldest is
unlinked, so the default reproduces the historical behaviour exactly: a
new recording replaces the previous one. ``0`` means unlimited, bounded
only by ``max_total_storage_mb``. ``3`` is a reasonable value for a fault
Expand Down
2 changes: 1 addition & 1 deletion docs/roadmap.rst
Original file line number Diff line number Diff line change
Expand Up @@ -123,7 +123,7 @@ and central aggregation.
**Key features:**

- [x] **Two-level filtering**: FaultReporter (local) + FaultManager (central)
- [x] **Multi-source aggregation**: Same fault code from multiple sources combined into single entry
- [x] **Per-source fault records**: Same fault code from multiple sources kept as one record each, addressed by (code, source)
- [x] **Persistent storage**: Fault state survives restarts
- [x] **REST API + SSE**: Real-time fault monitoring via HTTP
- [x] **Backwards compatibility**: Integration with ``diagnostic_updater``
Expand Down
2 changes: 1 addition & 1 deletion docs/tutorials/docker.rst
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ Images are available for all supported ROS 2 distributions:
Each image includes the gateway and all open-core packages:

- ``ros2_medkit_gateway`` - HTTP REST server
- ``ros2_medkit_fault_manager`` - Fault aggregation and management
- ``ros2_medkit_fault_manager`` - Fault record keeping and lifecycle management
- ``ros2_medkit_fault_reporter`` - Client library for fault reporting
- ``ros2_medkit_diagnostic_bridge`` - Bridges ``/diagnostics`` to fault manager
- ``ros2_medkit_serialization`` - Runtime JSON/ROS 2 serialization
Expand Down
6 changes: 3 additions & 3 deletions docs/tutorials/snapshots.rst
Original file line number Diff line number Diff line change
Expand Up @@ -118,11 +118,11 @@ Configure snapshot capture via fault manager parameters:
- Use background subscriptions (caches latest message)
* - ``snapshots.recapture_cooldown_sec``
- ``60.0``
- Minimum seconds between snapshot captures for the same fault code.
- Minimum seconds between snapshot captures for the same fault record.
Prevents snapshot storms when a fault is reported repeatedly. Set to 0 to disable.
* - ``snapshots.max_per_fault``
- ``10``
- Maximum number of snapshot rows stored per fault code. One confirmation
- Maximum number of snapshot rows stored per fault record. One confirmation
writes one row per configured topic, and those rows are evicted together:
past the limit the OLDEST capture set is dropped whole. A capture larger
than the cap is kept anyway rather than torn, since half a freeze frame
Expand Down Expand Up @@ -565,7 +565,7 @@ Rosbag Configuration Options
burst's bag at a time.
* - ``snapshots.rosbag.max_bags_per_fault``
- ``1``
- Recordings kept per fault code; ``0`` means unlimited. Past the cap the
- Recordings kept per fault record (``0`` means unlimited). Past the cap the
fault's oldest recording is dropped, and a bag is deleted only once no
fault still references it (a burst shares one recording). ``1`` is the
historical behaviour - each re-confirmation replaces the previous bag;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -132,8 +132,9 @@ class ActionStatusBridgeNode : public rclcpp::Node {
/// Get (creating on first use) the FaultReporter for an action. The reporter's
/// source_id is fixed when first created: the resolved server FQN if discovery
/// has settled, otherwise the action name as a fallback so the fault still fires
/// on time. It is never re-attributed afterwards (reporting_sources is
/// append-only and the per-entity scope filter is strict-AND).
/// on time. It is never re-attributed afterwards: the fault manager keeps one
/// record per (fault code, source), so a swapped source would open a second
/// record and leave the provisional one raised with nothing left to heal it.
ros2_medkit_fault_reporter::FaultReporter * reporter_for(const std::string & action_name);

/// Resolve the action server's node FQN from its status-topic publisher, for
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -476,11 +476,12 @@ ros2_medkit_fault_reporter::FaultReporter * ActionStatusBridgeNode::reporter_for
// so the fault still fires on time. Reporting the fault on time takes priority
// over entity attribution when discovery is slow.
//
// The source is NOT re-attributed later: reporting_sources is append-only on the
// manager side and the per-entity /faults scope filter is strict-AND, so a
// provisional action-name source cannot be swapped for the FQN afterwards.
// Correct attribution for the slow-discovery case is a separate concern (it
// needs a way to supersede a provisional source).
// The source is NOT re-attributed later: the fault manager keeps one record per
// (fault code, source), so swapping the provisional action-name source for the
// FQN would open a second record under the FQN and leave the provisional one
// raised, with no reporter left to send its PASSED. Correct attribution for the
// slow-discovery case is a separate concern (it needs a way to supersede a
// provisional source).
const std::string fqn = server_fqn_for_action(action_name);
const std::string & source_id = fqn.empty() ? action_name : fqn;
auto reporter = std::make_unique<ros2_medkit_fault_reporter::FaultReporter>(this->shared_from_this(), source_id);
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -283,7 +283,7 @@ TEST_F(ActionStatusBridgeTest, ReporterFor_StickyCreatedOnceNeverSwapped) {
ActionStatusBridgeTestAccess access(node.get());
// No publisher exists, so the server FQN is unresolved and the reporter falls
// back to the action name. It must be created once and reused, never swapped:
// reporting_sources is append-only, so a provisional source cannot be undone.
// a swapped source would open a second fault record and strand the first.
const void * first = access.reporter_identity("/nav");
EXPECT_NE(first, nullptr);
EXPECT_EQ(access.reporter_identity("/nav"), first);
Expand Down
Loading
Loading