Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
3100bf4
feat(gateway): carry the owning source on every per-record fault call
bburda Sep 14, 2026
27925fc
feat(gateway): resolve a per-entity fault route to one owned record
bburda Sep 14, 2026
7956e1c
feat(opcua): clear the record the plugin raised, not the code
bburda Sep 16, 2026
ee01095
fix(gateway): read a trigger rule's value from the app's own data source
bburda Sep 16, 2026
6dc24e7
test(integration): drive two owners of one fault code over HTTP
bburda Sep 16, 2026
1bdb0c3
feat(gateway): name the record each muted fault entry describes
bburda Sep 16, 2026
4fb44e8
fix(gateway): address the record a plugin clear and a bag URL mean
bburda Sep 16, 2026
aff0e80
fix(gateway): resolve a fault code over the records the entity lists
bburda Sep 24, 2026
ef0513d
fix(gateway): answer 503 when a plugin clear cannot read the store
bburda Sep 24, 2026
78c6ef3
feat(gateway): add FaultProvider::clear_fault_record, plugin API v8
bburda Sep 24, 2026
b84260e
fix(gateway): declare the recording download 409 and name real ids
bburda Sep 24, 2026
416cfb9
test(integration): make the per-code DELETE 409 check able to fail
bburda Sep 24, 2026
9148d4a
docs(gateway): state that the bulk fault DELETE skips muted records
bburda Sep 24, 2026
4ee53ae
fix(gateway): state the trigger engine's own rule in its 409
bburda Sep 24, 2026
4d32d2b
fix(opcua): keep the ClearFault reply callback off the plugin
bburda Sep 24, 2026
25a290c
style: split clauses joined by semicolons in comments and messages
bburda Sep 24, 2026
bd6be17
style: apply clang-format to the record resolution and plugin clear code
bburda Sep 24, 2026
d8448d5
fix(opcua): serve the status route only for the plugin's own component
bburda Sep 24, 2026
b9f4690
fix(gateway): prefer the records an entity lists when resolving a code
bburda Sep 24, 2026
7a37a34
fix(gateway): serve a recording only to an owner of a record it holds
bburda Sep 24, 2026
30c88dd
fix(gateway): coalesce fault stream events per record, not per code
bburda Sep 24, 2026
728aa6a
perf(opcua): read the ClearFault reply without copying it
bburda Sep 24, 2026
0bddea1
docs: drop code-keyed fault wording from the OPC UA plugin and tests
bburda Sep 24, 2026
0ba2da5
docs(gateway): state which owners a recording download accepts
bburda Sep 24, 2026
8fba96c
style: clear clang-tidy findings on added test lines
bburda Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 5 additions & 2 deletions docs/api/locking.rst
Original file line number Diff line number Diff line change
Expand Up @@ -232,8 +232,11 @@ The One Exception: Global Fault Clear
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

``DELETE /api/v1/faults`` reads ``X-Client-Id`` like the writes above but never
answers ``409``. It walks every fault, **skips** the ones whose reporting
entity is locked by another client, clears the rest, and answers ``204``.
answers ``409``. It walks every fault RECORD, **skips** the ones whose owning
entity is locked by another client, clears the rest one record at a time with
its owner, and answers ``204``. The owner is resolved the same way the
per-entity fault scope resolves it, so an external app or an external component
is found by its bare SOVD id and its lock is honoured here too.
Nothing on the response says which faults were skipped - the
``X-Medkit-Local-Only: true`` header that 204 also carries is set
unconditionally and reports that aggregated *peers* were not cleared, not that
Expand Down
16 changes: 11 additions & 5 deletions docs/api/messages.rst
Original file line number Diff line number Diff line change
Expand Up @@ -255,13 +255,19 @@ When ``skip_correlation_auto_clear`` is ``false`` (default), clearing a
root-cause fault also clears every symptom that the correlation engine
attributes to it via ``auto_clear_with_root`` rules; the cleared
symptom codes are returned in ``auto_cleared_codes``. When ``true``,
only the requested ``fault_code`` is cleared and ``auto_cleared_codes``
is empty. The gateway's per-entity ``DELETE
only the requested record is cleared and ``auto_cleared_codes`` is
empty. The gateway's per-entity ``DELETE
/{entity-path}/faults/{fault_code}`` route sets this to ``true`` so
that an operator with access to one entity cannot cascade-clear
correlated symptoms reported by apps in other entities. The global
``DELETE /api/v1/faults/{fault_code}`` route leaves it ``false`` so
cluster-wide clearing still works.
correlated symptoms reported by apps in other entities, and its
per-entity ``DELETE /{entity-path}/faults`` does the same for every
record it clears. The global ``DELETE /api/v1/faults`` route leaves it
``false`` so cluster-wide clearing still works. There is no
``DELETE /api/v1/faults/{fault_code}`` route: a bare fault code is not
an address, because it names as many records as there are sources
reporting it. A client that wants the cascade for one record calls
``ClearFault`` directly with that record's ``fault_code`` and
``source_id``.

.. note::

Expand Down
120 changes: 100 additions & 20 deletions docs/api/rest.rst
Original file line number Diff line number Diff line change
Expand Up @@ -1374,20 +1374,54 @@ Query and manage faults.

.. note::

**Per-entity fault scope (``/{entity-path}/faults`` routes).** The gateway keys
faults by ``fault_code`` only, and a fault's ``reporting_sources`` set is the
union of every app that has reported that code. Per-entity routes apply a
strict all-sources scope check: a fault is in scope for an entity iff **every**
entry in ``reporting_sources`` is an app owned by that entity (exact FQN
match, or strict path-child).

This means a ``fault_code`` reported by apps in two different entities
(for example ``SENSOR_TIMEOUT`` reported by both the lidar and the
temperature sensor app) is **not** visible or clearable through either
entity's per-entity routes - per-fault routes return ``404``, collection
responses omit it, and per-entity ``DELETE`` skips it. To see, list, or
clear such shared faults use the global ``GET /api/v1/faults`` /
``DELETE /api/v1/faults`` routes.
**Per-entity fault scope (``/{entity-path}/faults`` routes).** A fault record
is the pair (``fault_code``, reporting source). The source is the ``source_id``
the reporter used, it owns the record, and it is the single entry of the
record's ``reporting_sources``. Two apps reporting one ``fault_code`` are two
records, each with its own status, occurrence count and timestamps, and each
cleared on its own. A record is in scope for an entity when its owner is one
of the entity's own reporting sources (an exact FQN match or a strict
path-child, and for an external app or an external component its bare SOVD
id).

``SENSOR_TIMEOUT`` reported by both the lidar app and the temperature sensor
app therefore appears on each app's own ``/faults`` page as that app's record,
and on the component hosting both as two items with distinct ``source_id``.

A ``fault_code`` in a URL names a record only when it resolves to exactly one
record in the entity's scope. The candidates come in three tiers, and the
first tier holding any record of the code decides:

1. the records the entity's fault list shows without a ``status`` parameter,
which are the ``PREFAILED`` and ``CONFIRMED`` records that are not muted
2. the muted records of those statuses, which that list leaves out only
because the correlation engine muted them
3. the records the list shows only for ``status=cleared``, ``status=healed``
or ``status=all``: ``CLEARED``, ``HEALED`` and ``PREPASSED``

So a muted symptom, or a record its source cleared, stays addressable by its
code while nothing ranks above it, and it never makes the record a client read
off the list ambiguous. Once one of two sources has cleared its record, the
code names the other one. ``GET`` and ``DELETE`` on
``/{entity-path}/faults/{fault_code}`` and the recording download
``GET /{entity-path}/bulk-data/rosbags/{fault_code}`` all resolve this way.
One candidate, and the route acts on it with its owner. Several, and it
answers ``409`` with vendor error code ``x-medkit-ambiguous-fault`` and
``parameters.owners`` naming them, rather than acting on whichever record the
store listed first. Address one of them through the route of the app that
owns it, ``/apps/{app_id}/faults/{fault_code}``. None, and it answers ``404``.
Per-entity ``DELETE /{entity-path}/faults`` clears the records the entity's
fault list shows, each individually with its own owner. Like that list, it
leaves muted records alone. A muted record is cleared by its own per-code
``DELETE /{entity-path}/faults/{fault_code}``, which resolves it as above.

Two nearby keys carry the record's owner and they are deliberately not the
same name. A flat fault item's top-level ``source_id`` is the owner, and the
detail response's ``x-medkit.owner`` is the same value. A fault LIST's
``x-medkit.source_id`` is something else entirely: the addressed entity's own
namespace path, which is empty for an external app. Read ``source_id`` on an
item, ``owner`` on a detail, and never the list-level ``source_id`` as an
owner.

``GET /api/v1/faults``
List all faults across the system.
Expand Down Expand Up @@ -1483,6 +1517,7 @@ Query and manage faults.
"x-medkit": {
"occurrence_count": 3,
"reporting_sources": ["/powertrain/motor_controller"],
"owner": "/powertrain/motor_controller",
"severity_label": "ERROR"
}
}
Expand Down Expand Up @@ -1532,16 +1567,25 @@ Query and manage faults.
- **400:** ``fault_code`` empty or longer than 256 characters
- **404:** Fault not found, reported by an app outside this entity's scope,
or declined by the fault manager
- **409:** ``x-medkit-ambiguous-fault``, the code names several records in
this entity's scope
- **503:** Fault manager unavailable

``DELETE /api/v1/components/{id}/faults/{fault_code}``
Clear a fault.

- **200:** Fault cleared by the plugin that serves this entity's faults, with
the plugin's acknowledgement as the body
- **204:** Fault cleared
- **400:** ``fault_code`` empty or longer than 256 characters
- **404:** Fault not found, reported by an app outside this entity's scope,
or declined by the fault manager
- **503:** Fault manager unavailable
- **409:** ``x-medkit-ambiguous-fault``, the code names several records in
this entity's scope, or the entity is locked by another client
- **503:** Fault manager unavailable. This holds for an entity whose faults a
plugin serves as well: the gateway reads the fault manager to learn which
record the code names before it asks the plugin, and when that read fails
it does not ask the plugin at all.

.. note::

Expand All @@ -1561,7 +1605,9 @@ Query and manage faults.
Clear all faults for an entity.

Accepts the optional ``?status=`` query parameter (same values as ``GET /faults``).
Without it, clears pending and confirmed faults.
Without it, clears pending and confirmed faults. Muted records are not cleared,
because the entity's fault list does not show them. Clear one with its per-code
``DELETE``.

- **204:** Faults cleared (or none to clear)
- **400:** Invalid status parameter
Expand Down Expand Up @@ -1822,6 +1868,15 @@ and OpenAPI 3.1 has no way to say "bytes" - ``format: binary`` was an OpenAPI
3.0 idiom that 3.1 dropped when it aligned with JSON Schema 2020-12. A
schema-free media type entry is the accurate description.

**Which recordings an entity serves.** A recording belongs to the fault records
it is attached to, each a (``fault_code``, owner) pair, and a burst attaches
several. An entity serves a recording when the owner of one of those records is
in the entity's fault scope, the scope its fault list uses. A fault code the
entity also owns a record of is not enough:
two apps reporting one code each download their own recording and get ``404`` on
the other's, whether or not either record has been cleared, while the component
hosting both serves both.

**Range requests.** A request carrying a ``Range`` header is answered with
**206 Partial Content** and a ``Content-Range: bytes <start>-<end>/<total>``
header instead of ``200``; the body is the requested slice. Several ranges in
Expand All @@ -1842,7 +1897,16 @@ declared on the 206 only - the 200 can never carry it.

- **200 OK**: File content
- **206 Partial Content**: The byte range requested via ``Range``, with ``Content-Range``
- **404 Not Found**: Entity, category, or bulk-data ID not found
- **404 Not Found**: Entity, category, or bulk-data ID not found, or a
recording none of this entity's records is attached to
- **409 Conflict**: ``x-medkit-ambiguous-fault``. A ``rosbags`` URL whose last
segment is a fault code rather than a recording id resolves that code to one
record the way the per-code fault routes do (see the fault record note under
Faults Endpoints), and several candidate records in the entity's scope name
none of them. ``parameters.owners`` names the owners. Address the recording by
its own id instead: the ``rosbags`` listing of this entity carries it as each
descriptor's ``id``, and a fault detail links it as
``environment_data.snapshots[].bulk_data_uri``.
- **416 Range Not Satisfiable**: The ``Range`` header could not be parsed. Not
specific to this endpoint - see :ref:`rest-range-rejection`.

Expand Down Expand Up @@ -2807,6 +2871,14 @@ plugin is loaded. Without it the routes stay mounted and answer ``501``
use, so a client can tell "this build has no threshold engine" apart from "no
such app or rule".

A rule reads its value from whichever source the owning plugin serves the app's
``/data`` through: the plugin's data provider first, its own vendor data route
only when it exposes no provider. Rule evaluation and the create-time data-point
check use the same resolution in the same order, so a point that validates on
create is one the engine can read. A rule the engine cannot read holds its state
rather than firing on a stale value, which is also what it does while the app's
link reports itself down.

``GET /api/v1/apps/{app_id}/fault-triggers``
List the app's rules. The owning app is the one in the path; it is not
repeated in the item, and neither is the engine's internal cross latch.
Expand Down Expand Up @@ -2838,8 +2910,9 @@ such app or rule".
``data_name`` the app does not expose (when enumerable); ``404``
(``entity-not-found``) when the app itself was never discovered; ``409``
(``precondition-not-fulfilled``) when the ``fault_code`` is already used by
another rule - fault codes are global to the fault store, so two rules
sharing one would fight over the same fault.
another rule on ANY app - the engine keeps one rule per code across every
app, so a code another app's rule already claims is refused. The error names
the rule and the app that holds it.

``DELETE /api/v1/apps/{app_id}/fault-triggers/{trigger_id}``
Remove a rule (``204``). A fault currently asserted by the rule is cleared;
Expand Down Expand Up @@ -3394,6 +3467,13 @@ Vendor-specific ``x-medkit-*`` codes are enveloped: the response carries
* - ``vendor-error``
- varies
- A vendor-specific failure; read ``vendor_code`` for the real code
* - ``x-medkit-ambiguous-fault``
- 409
- The ``fault_code`` in the URL addresses several fault records inside the
addressed entity, because several of its reporting sources report that
code and each is its own record. ``parameters.owners`` names them.
Address one of them through the route of the app that owns it,
``/apps/{app_id}/faults/{fault_code}``.
* - ``x-medkit-plugin-error``
- 400-599
- Plugin provider returned an error. Status varies by plugin. Message truncated to 512 chars.
Expand Down Expand Up @@ -3829,7 +3909,7 @@ Other extensions beyond SOVD:
- ``DELETE /faults`` - Clear all faults globally
- ``GET /faults/stream`` - SSE real-time fault notifications. Each event payload carries an
optional ``x-medkit`` SOVD payload-extension object with ``entity_type`` and ``entity_id``
fields when the gateway can resolve the fault's first reporting source back to an entity,
fields when the gateway can resolve the fault record's reporting source (its owner) to an entity,
so consumers can hit ``/{entity_type}/{entity_id}/bulk-data/rosbags/{fault_code}`` directly
without enumerating entities - that address serves the fault's newest recording. To reach an
older one, list ``/bulk-data/rosbags`` and use the descriptor ``id``. Resolution is snapshotted at event arrival; the entire
Expand Down
16 changes: 9 additions & 7 deletions docs/design/ros2_medkit_fault_detection/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -62,10 +62,12 @@ Transition tracking and global uniqueness
the evaluator. It returns only the signals whose active state changed since the
last call, which a plugin forwards to the fault manager as report / clear.

The tracker is keyed by ``fault_code`` alone, matching the fault manager, which
also keys and clears faults by code alone. A single tracker may therefore be
shared across many points only if every ``fault_code`` is globally unique; two
points emitting the same code would alternately raise and clear it each cycle.
Consumers that share one tracker (for example the OPC UA poller across all
node-map entries and event alarms) must enforce that uniqueness at
config-load time and reject a colliding configuration before anything runs.
The tracker is keyed by ``fault_code`` alone. That is a property of the tracker,
not of the fault manager, which identifies a record by ``fault_code`` and the
reporting source that owns it. A single tracker may therefore be shared across
many points only if every ``fault_code`` is unique within it. Two points emitting
the same code would alternately raise and clear it each cycle, inside the tracker
and before any report is sent. Consumers that share one tracker (for example the
OPC UA poller across all node-map entries and event alarms) must enforce that
uniqueness at config-load time and reject a colliding configuration before
anything runs.
7 changes: 5 additions & 2 deletions docs/glossary.rst
Original file line number Diff line number Diff line change
Expand Up @@ -55,8 +55,11 @@ This glossary defines key terms used throughout ros2_medkit documentation.
can be polled for status or cancelled.

Fault
An error condition reported by a ROS 2 node to the fault manager.
Faults have a code, severity, message, and timestamp.
An error condition reported to the fault manager. One fault record is the
pair (fault code, reporting source), the source being the ``source_id``
the reporter used. Two sources reporting one code are two records, each
with its own severity, message, timestamps and status, and each cleared on
its own.

See: :doc:`design/ros2_medkit_fault_reporter/index`

Expand Down
3 changes: 1 addition & 2 deletions docs/requirements/specs/faults.rst
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ Faults
- ``item``: Fault details with SOVD-compliant ``status`` object (aggregatedStatus, testFailed, confirmedDTC, pendingDTC)
- ``environment_data``: Extended data records (timestamps) and snapshots array
- ``environment_data.snapshots[]``: Array of freeze_frame (topic data) and rosbag (bulk-data reference) entries
- ``x-medkit``: Extension fields (occurrence_count, reporting_sources, severity_label)
- ``x-medkit``: Extension fields (occurrence_count, reporting_sources, owner, severity_label)

.. req:: DELETE /{entity}/faults
:id: REQ_INTEROP_014
Expand Down Expand Up @@ -59,4 +59,3 @@ Faults
endpoints on the fault's reporting entity. For faults already confirmed when the
capturing component starts (e.g. the gateway restarts while a fault is standing), an
equivalent snapshot shall be captured at startup and marked with its capture origin.

29 changes: 17 additions & 12 deletions docs/tutorials/fault-correlation.rst
Original file line number Diff line number Diff line change
Expand Up @@ -339,7 +339,8 @@ Response always includes:
"fault_code": "MOTOR_COMM_001",
"root_cause_code": "ESTOP_001",
"rule_id": "estop_cascade",
"delay_ms": 150
"delay_ms": 150,
"source_id": "/powertrain/motor_controller"
}
]
}
Expand Down Expand Up @@ -385,17 +386,21 @@ Response always includes:

.. note::

**Per-entity DELETE opts out of the cascade.** The same fault cleared
via ``DELETE /api/v1/{entity-path}/faults/ESTOP_001`` clears only
``ESTOP_001`` itself - ``auto_cleared_codes`` will be empty in the
response. The gateway sets ``ClearFault.srv``'s
``skip_correlation_auto_clear`` to ``true`` on per-entity routes so
that an operator with access to one entity cannot cascade-clear
correlated symptom faults reported by apps in other entities. Use
the global ``DELETE /api/v1/faults/{fault_code}`` route when you do
want the correlation cascade. Direct ``ros2 service call`` clients
can choose explicitly by setting ``skip_correlation_auto_clear`` in
the request body (see ``ros2_medkit_msgs`` ``ClearFault.srv``).
**Per-entity DELETE opts out of the cascade.** The same record cleared
via ``DELETE /api/v1/{entity-path}/faults/ESTOP_001`` clears only that
record - ``auto_cleared_codes`` will be empty in the response. The
gateway sets ``ClearFault.srv``'s ``skip_correlation_auto_clear`` to
``true`` on per-entity routes so that an operator with access to one
entity cannot cascade-clear correlated symptom faults reported by apps
in other entities.

There is no global ``DELETE /api/v1/faults/{fault_code}`` route over
HTTP: a bare fault code names as many records as there are sources
reporting it, so it is not an address. ``DELETE /api/v1/faults``
clears every record and takes no code. To get the cascade for one
record, call the service directly and set ``skip_correlation_auto_clear``
to ``false`` alongside that record's ``fault_code`` and ``source_id``
(see ``ros2_medkit_msgs`` ``ClearFault.srv``).

Example: Complete Configuration
-------------------------------
Expand Down
Loading