diff --git a/docs/api/locking.rst b/docs/api/locking.rst index 50c382633..8203085b7 100644 --- a/docs/api/locking.rst +++ b/docs/api/locking.rst @@ -232,8 +232,11 @@ The One Exception: Global Fault Clear ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ``DELETE /api/v1/faults`` reads ``X-Client-Id`` like the writes above but never -answers ``409``. It walks every fault, **skips** the ones whose reporting -entity is locked by another client, clears the rest, and answers ``204``. +answers ``409``. It walks every fault RECORD, **skips** the ones whose owning +entity is locked by another client, clears the rest one record at a time with +its owner, and answers ``204``. The owner is resolved the same way the +per-entity fault scope resolves it, so an external app or an external component +is found by its bare SOVD id and its lock is honoured here too. Nothing on the response says which faults were skipped - the ``X-Medkit-Local-Only: true`` header that 204 also carries is set unconditionally and reports that aggregated *peers* were not cleared, not that diff --git a/docs/api/messages.rst b/docs/api/messages.rst index a4b6b4860..29298c429 100644 --- a/docs/api/messages.rst +++ b/docs/api/messages.rst @@ -255,13 +255,19 @@ When ``skip_correlation_auto_clear`` is ``false`` (default), clearing a root-cause fault also clears every symptom that the correlation engine attributes to it via ``auto_clear_with_root`` rules; the cleared symptom codes are returned in ``auto_cleared_codes``. When ``true``, -only the requested ``fault_code`` is cleared and ``auto_cleared_codes`` -is empty. The gateway's per-entity ``DELETE +only the requested record is cleared and ``auto_cleared_codes`` is +empty. The gateway's per-entity ``DELETE /{entity-path}/faults/{fault_code}`` route sets this to ``true`` so that an operator with access to one entity cannot cascade-clear -correlated symptoms reported by apps in other entities. The global -``DELETE /api/v1/faults/{fault_code}`` route leaves it ``false`` so -cluster-wide clearing still works. +correlated symptoms reported by apps in other entities, and its +per-entity ``DELETE /{entity-path}/faults`` does the same for every +record it clears. The global ``DELETE /api/v1/faults`` route leaves it +``false`` so cluster-wide clearing still works. There is no +``DELETE /api/v1/faults/{fault_code}`` route: a bare fault code is not +an address, because it names as many records as there are sources +reporting it. A client that wants the cascade for one record calls +``ClearFault`` directly with that record's ``fault_code`` and +``source_id``. .. note:: diff --git a/docs/api/rest.rst b/docs/api/rest.rst index 5f55fb862..ffa3fa713 100644 --- a/docs/api/rest.rst +++ b/docs/api/rest.rst @@ -1374,20 +1374,54 @@ Query and manage faults. .. note:: - **Per-entity fault scope (``/{entity-path}/faults`` routes).** The gateway keys - faults by ``fault_code`` only, and a fault's ``reporting_sources`` set is the - union of every app that has reported that code. Per-entity routes apply a - strict all-sources scope check: a fault is in scope for an entity iff **every** - entry in ``reporting_sources`` is an app owned by that entity (exact FQN - match, or strict path-child). - - This means a ``fault_code`` reported by apps in two different entities - (for example ``SENSOR_TIMEOUT`` reported by both the lidar and the - temperature sensor app) is **not** visible or clearable through either - entity's per-entity routes - per-fault routes return ``404``, collection - responses omit it, and per-entity ``DELETE`` skips it. To see, list, or - clear such shared faults use the global ``GET /api/v1/faults`` / - ``DELETE /api/v1/faults`` routes. + **Per-entity fault scope (``/{entity-path}/faults`` routes).** A fault record + is the pair (``fault_code``, reporting source). The source is the ``source_id`` + the reporter used, it owns the record, and it is the single entry of the + record's ``reporting_sources``. Two apps reporting one ``fault_code`` are two + records, each with its own status, occurrence count and timestamps, and each + cleared on its own. A record is in scope for an entity when its owner is one + of the entity's own reporting sources (an exact FQN match or a strict + path-child, and for an external app or an external component its bare SOVD + id). + + ``SENSOR_TIMEOUT`` reported by both the lidar app and the temperature sensor + app therefore appears on each app's own ``/faults`` page as that app's record, + and on the component hosting both as two items with distinct ``source_id``. + + A ``fault_code`` in a URL names a record only when it resolves to exactly one + record in the entity's scope. The candidates come in three tiers, and the + first tier holding any record of the code decides: + + 1. the records the entity's fault list shows without a ``status`` parameter, + which are the ``PREFAILED`` and ``CONFIRMED`` records that are not muted + 2. the muted records of those statuses, which that list leaves out only + because the correlation engine muted them + 3. the records the list shows only for ``status=cleared``, ``status=healed`` + or ``status=all``: ``CLEARED``, ``HEALED`` and ``PREPASSED`` + + So a muted symptom, or a record its source cleared, stays addressable by its + code while nothing ranks above it, and it never makes the record a client read + off the list ambiguous. Once one of two sources has cleared its record, the + code names the other one. ``GET`` and ``DELETE`` on + ``/{entity-path}/faults/{fault_code}`` and the recording download + ``GET /{entity-path}/bulk-data/rosbags/{fault_code}`` all resolve this way. + One candidate, and the route acts on it with its owner. Several, and it + answers ``409`` with vendor error code ``x-medkit-ambiguous-fault`` and + ``parameters.owners`` naming them, rather than acting on whichever record the + store listed first. Address one of them through the route of the app that + owns it, ``/apps/{app_id}/faults/{fault_code}``. None, and it answers ``404``. + Per-entity ``DELETE /{entity-path}/faults`` clears the records the entity's + fault list shows, each individually with its own owner. Like that list, it + leaves muted records alone. A muted record is cleared by its own per-code + ``DELETE /{entity-path}/faults/{fault_code}``, which resolves it as above. + + Two nearby keys carry the record's owner and they are deliberately not the + same name. A flat fault item's top-level ``source_id`` is the owner, and the + detail response's ``x-medkit.owner`` is the same value. A fault LIST's + ``x-medkit.source_id`` is something else entirely: the addressed entity's own + namespace path, which is empty for an external app. Read ``source_id`` on an + item, ``owner`` on a detail, and never the list-level ``source_id`` as an + owner. ``GET /api/v1/faults`` List all faults across the system. @@ -1483,6 +1517,7 @@ Query and manage faults. "x-medkit": { "occurrence_count": 3, "reporting_sources": ["/powertrain/motor_controller"], + "owner": "/powertrain/motor_controller", "severity_label": "ERROR" } } @@ -1532,16 +1567,25 @@ Query and manage faults. - **400:** ``fault_code`` empty or longer than 256 characters - **404:** Fault not found, reported by an app outside this entity's scope, or declined by the fault manager + - **409:** ``x-medkit-ambiguous-fault``, the code names several records in + this entity's scope - **503:** Fault manager unavailable ``DELETE /api/v1/components/{id}/faults/{fault_code}`` Clear a fault. + - **200:** Fault cleared by the plugin that serves this entity's faults, with + the plugin's acknowledgement as the body - **204:** Fault cleared - **400:** ``fault_code`` empty or longer than 256 characters - **404:** Fault not found, reported by an app outside this entity's scope, or declined by the fault manager - - **503:** Fault manager unavailable + - **409:** ``x-medkit-ambiguous-fault``, the code names several records in + this entity's scope, or the entity is locked by another client + - **503:** Fault manager unavailable. This holds for an entity whose faults a + plugin serves as well: the gateway reads the fault manager to learn which + record the code names before it asks the plugin, and when that read fails + it does not ask the plugin at all. .. note:: @@ -1561,7 +1605,9 @@ Query and manage faults. Clear all faults for an entity. Accepts the optional ``?status=`` query parameter (same values as ``GET /faults``). - Without it, clears pending and confirmed faults. + Without it, clears pending and confirmed faults. Muted records are not cleared, + because the entity's fault list does not show them. Clear one with its per-code + ``DELETE``. - **204:** Faults cleared (or none to clear) - **400:** Invalid status parameter @@ -1822,6 +1868,15 @@ and OpenAPI 3.1 has no way to say "bytes" - ``format: binary`` was an OpenAPI 3.0 idiom that 3.1 dropped when it aligned with JSON Schema 2020-12. A schema-free media type entry is the accurate description. +**Which recordings an entity serves.** A recording belongs to the fault records +it is attached to, each a (``fault_code``, owner) pair, and a burst attaches +several. An entity serves a recording when the owner of one of those records is +in the entity's fault scope, the scope its fault list uses. A fault code the +entity also owns a record of is not enough: +two apps reporting one code each download their own recording and get ``404`` on +the other's, whether or not either record has been cleared, while the component +hosting both serves both. + **Range requests.** A request carrying a ``Range`` header is answered with **206 Partial Content** and a ``Content-Range: bytes -/`` header instead of ``200``; the body is the requested slice. Several ranges in @@ -1842,7 +1897,16 @@ declared on the 206 only - the 200 can never carry it. - **200 OK**: File content - **206 Partial Content**: The byte range requested via ``Range``, with ``Content-Range`` -- **404 Not Found**: Entity, category, or bulk-data ID not found +- **404 Not Found**: Entity, category, or bulk-data ID not found, or a + recording none of this entity's records is attached to +- **409 Conflict**: ``x-medkit-ambiguous-fault``. A ``rosbags`` URL whose last + segment is a fault code rather than a recording id resolves that code to one + record the way the per-code fault routes do (see the fault record note under + Faults Endpoints), and several candidate records in the entity's scope name + none of them. ``parameters.owners`` names the owners. Address the recording by + its own id instead: the ``rosbags`` listing of this entity carries it as each + descriptor's ``id``, and a fault detail links it as + ``environment_data.snapshots[].bulk_data_uri``. - **416 Range Not Satisfiable**: The ``Range`` header could not be parsed. Not specific to this endpoint - see :ref:`rest-range-rejection`. @@ -2807,6 +2871,14 @@ plugin is loaded. Without it the routes stay mounted and answer ``501`` use, so a client can tell "this build has no threshold engine" apart from "no such app or rule". +A rule reads its value from whichever source the owning plugin serves the app's +``/data`` through: the plugin's data provider first, its own vendor data route +only when it exposes no provider. Rule evaluation and the create-time data-point +check use the same resolution in the same order, so a point that validates on +create is one the engine can read. A rule the engine cannot read holds its state +rather than firing on a stale value, which is also what it does while the app's +link reports itself down. + ``GET /api/v1/apps/{app_id}/fault-triggers`` List the app's rules. The owning app is the one in the path; it is not repeated in the item, and neither is the engine's internal cross latch. @@ -2838,8 +2910,9 @@ such app or rule". ``data_name`` the app does not expose (when enumerable); ``404`` (``entity-not-found``) when the app itself was never discovered; ``409`` (``precondition-not-fulfilled``) when the ``fault_code`` is already used by - another rule - fault codes are global to the fault store, so two rules - sharing one would fight over the same fault. + another rule on ANY app - the engine keeps one rule per code across every + app, so a code another app's rule already claims is refused. The error names + the rule and the app that holds it. ``DELETE /api/v1/apps/{app_id}/fault-triggers/{trigger_id}`` Remove a rule (``204``). A fault currently asserted by the rule is cleared; @@ -3394,6 +3467,13 @@ Vendor-specific ``x-medkit-*`` codes are enveloped: the response carries * - ``vendor-error`` - varies - A vendor-specific failure; read ``vendor_code`` for the real code + * - ``x-medkit-ambiguous-fault`` + - 409 + - The ``fault_code`` in the URL addresses several fault records inside the + addressed entity, because several of its reporting sources report that + code and each is its own record. ``parameters.owners`` names them. + Address one of them through the route of the app that owns it, + ``/apps/{app_id}/faults/{fault_code}``. * - ``x-medkit-plugin-error`` - 400-599 - Plugin provider returned an error. Status varies by plugin. Message truncated to 512 chars. @@ -3829,7 +3909,7 @@ Other extensions beyond SOVD: - ``DELETE /faults`` - Clear all faults globally - ``GET /faults/stream`` - SSE real-time fault notifications. Each event payload carries an optional ``x-medkit`` SOVD payload-extension object with ``entity_type`` and ``entity_id`` - fields when the gateway can resolve the fault's first reporting source back to an entity, + fields when the gateway can resolve the fault record's reporting source (its owner) to an entity, so consumers can hit ``/{entity_type}/{entity_id}/bulk-data/rosbags/{fault_code}`` directly without enumerating entities - that address serves the fault's newest recording. To reach an older one, list ``/bulk-data/rosbags`` and use the descriptor ``id``. Resolution is snapshotted at event arrival; the entire diff --git a/docs/design/ros2_medkit_fault_detection/index.rst b/docs/design/ros2_medkit_fault_detection/index.rst index 4a79c1761..a4e11f80c 100644 --- a/docs/design/ros2_medkit_fault_detection/index.rst +++ b/docs/design/ros2_medkit_fault_detection/index.rst @@ -62,10 +62,12 @@ Transition tracking and global uniqueness the evaluator. It returns only the signals whose active state changed since the last call, which a plugin forwards to the fault manager as report / clear. -The tracker is keyed by ``fault_code`` alone, matching the fault manager, which -also keys and clears faults by code alone. A single tracker may therefore be -shared across many points only if every ``fault_code`` is globally unique; two -points emitting the same code would alternately raise and clear it each cycle. -Consumers that share one tracker (for example the OPC UA poller across all -node-map entries and event alarms) must enforce that uniqueness at -config-load time and reject a colliding configuration before anything runs. +The tracker is keyed by ``fault_code`` alone. That is a property of the tracker, +not of the fault manager, which identifies a record by ``fault_code`` and the +reporting source that owns it. A single tracker may therefore be shared across +many points only if every ``fault_code`` is unique within it. Two points emitting +the same code would alternately raise and clear it each cycle, inside the tracker +and before any report is sent. Consumers that share one tracker (for example the +OPC UA poller across all node-map entries and event alarms) must enforce that +uniqueness at config-load time and reject a colliding configuration before +anything runs. diff --git a/docs/glossary.rst b/docs/glossary.rst index 6e816d98c..9cc02df97 100644 --- a/docs/glossary.rst +++ b/docs/glossary.rst @@ -55,8 +55,11 @@ This glossary defines key terms used throughout ros2_medkit documentation. can be polled for status or cancelled. Fault - An error condition reported by a ROS 2 node to the fault manager. - Faults have a code, severity, message, and timestamp. + An error condition reported to the fault manager. One fault record is the + pair (fault code, reporting source), the source being the ``source_id`` + the reporter used. Two sources reporting one code are two records, each + with its own severity, message, timestamps and status, and each cleared on + its own. See: :doc:`design/ros2_medkit_fault_reporter/index` diff --git a/docs/requirements/specs/faults.rst b/docs/requirements/specs/faults.rst index f3434a3f6..d4ee31ec3 100644 --- a/docs/requirements/specs/faults.rst +++ b/docs/requirements/specs/faults.rst @@ -21,7 +21,7 @@ Faults - ``item``: Fault details with SOVD-compliant ``status`` object (aggregatedStatus, testFailed, confirmedDTC, pendingDTC) - ``environment_data``: Extended data records (timestamps) and snapshots array - ``environment_data.snapshots[]``: Array of freeze_frame (topic data) and rosbag (bulk-data reference) entries - - ``x-medkit``: Extension fields (occurrence_count, reporting_sources, severity_label) + - ``x-medkit``: Extension fields (occurrence_count, reporting_sources, owner, severity_label) .. req:: DELETE /{entity}/faults :id: REQ_INTEROP_014 @@ -59,4 +59,3 @@ Faults endpoints on the fault's reporting entity. For faults already confirmed when the capturing component starts (e.g. the gateway restarts while a fault is standing), an equivalent snapshot shall be captured at startup and marked with its capture origin. - diff --git a/docs/tutorials/fault-correlation.rst b/docs/tutorials/fault-correlation.rst index 07357780c..354e60f64 100644 --- a/docs/tutorials/fault-correlation.rst +++ b/docs/tutorials/fault-correlation.rst @@ -339,7 +339,8 @@ Response always includes: "fault_code": "MOTOR_COMM_001", "root_cause_code": "ESTOP_001", "rule_id": "estop_cascade", - "delay_ms": 150 + "delay_ms": 150, + "source_id": "/powertrain/motor_controller" } ] } @@ -385,17 +386,21 @@ Response always includes: .. note:: - **Per-entity DELETE opts out of the cascade.** The same fault cleared - via ``DELETE /api/v1/{entity-path}/faults/ESTOP_001`` clears only - ``ESTOP_001`` itself - ``auto_cleared_codes`` will be empty in the - response. The gateway sets ``ClearFault.srv``'s - ``skip_correlation_auto_clear`` to ``true`` on per-entity routes so - that an operator with access to one entity cannot cascade-clear - correlated symptom faults reported by apps in other entities. Use - the global ``DELETE /api/v1/faults/{fault_code}`` route when you do - want the correlation cascade. Direct ``ros2 service call`` clients - can choose explicitly by setting ``skip_correlation_auto_clear`` in - the request body (see ``ros2_medkit_msgs`` ``ClearFault.srv``). + **Per-entity DELETE opts out of the cascade.** The same record cleared + via ``DELETE /api/v1/{entity-path}/faults/ESTOP_001`` clears only that + record - ``auto_cleared_codes`` will be empty in the response. The + gateway sets ``ClearFault.srv``'s ``skip_correlation_auto_clear`` to + ``true`` on per-entity routes so that an operator with access to one + entity cannot cascade-clear correlated symptom faults reported by apps + in other entities. + + There is no global ``DELETE /api/v1/faults/{fault_code}`` route over + HTTP: a bare fault code names as many records as there are sources + reporting it, so it is not an address. ``DELETE /api/v1/faults`` + clears every record and takes no code. To get the cascade for one + record, call the service directly and set ``skip_correlation_auto_clear`` + to ``false`` alongside that record's ``fault_code`` and ``source_id`` + (see ``ros2_medkit_msgs`` ``ClearFault.srv``). Example: Complete Configuration ------------------------------- diff --git a/docs/tutorials/plugin-system.rst b/docs/tutorials/plugin-system.rst index 622b50e17..68899aca4 100644 --- a/docs/tutorials/plugin-system.rst +++ b/docs/tutorials/plugin-system.rst @@ -26,7 +26,14 @@ Plugins implement the ``GatewayPlugin`` C++ base class plus one or more typed pr - **OperationProvider** - per-entity operation backend (list operations, execute). Uses the same per-entity routing model as DataProvider. - **FaultProvider** - per-entity fault backend (list faults, get fault details, clear - faults). Uses the same per-entity routing model as DataProvider. + faults). Uses the same per-entity routing model as DataProvider. A fault record is + the pair (``fault_code``, owning reporting source), and the gateway clears one record + at a time through ``clear_fault_record(entity_id, fault_code, owner)``, where + ``owner`` is the source the gateway resolved in the entity's fault scope (for a + component, one of its hosted apps). Its default implementation calls + ``clear_fault(entity_id, fault_code)``, so a provider that overrides only the + two-argument clear keeps compiling and keeps receiving clears by code. Override + ``clear_fault_record`` to clear exactly the record the owner names. A single plugin can implement multiple provider interfaces. For example, a "systemd" plugin could provide both introspection (discover systemd units) and updates (manage service restarts). @@ -1019,8 +1026,17 @@ Plugins export ``plugin_api_version()`` which must return the gateway's ``PLUGIN If the version does not match, the plugin is rejected with a clear error message suggesting a rebuild against matching gateway headers. -The current API version is **5**. It is incremented when the ``PluginContext`` vtable changes -or breaking changes are made to ``GatewayPlugin`` or provider interfaces. +The current API version is **8**. It is incremented when the ``PluginContext`` vtable changes +or breaking changes are made to ``GatewayPlugin`` or provider interfaces. A new virtual method +counts even when it keeps plugin source compiling, because it changes the vtable a pre-built +plugin was compiled against. Version 8 added ``FaultProvider::clear_fault_record``: its +default implementation calls the unchanged ``clear_fault``, so v7 plugin source compiles +against v8 headers as it is, but a plugin binary built for v7 is rejected and has to be +rebuilt. Version 8 also gives ``GatewayPlugin`` a protected ``log_sink()``, a copy of the +plugin's log sink for work that can outlive the plugin (an asynchronous service reply, for +example, should log through that copy rather than capture ``this``), and makes +``set_logger()`` protected so a plugin hosted without the gateway, in a unit test for +example, can wire a sink of its own. Build Requirements ------------------ diff --git a/docs/tutorials/snapshots.rst b/docs/tutorials/snapshots.rst index 85af31467..a105830db 100644 --- a/docs/tutorials/snapshots.rst +++ b/docs/tutorials/snapshots.rst @@ -310,7 +310,8 @@ Snapshots are included inline in the fault response as ``environment_data``: }, "x-medkit": { "occurrence_count": 3, - "reporting_sources": ["/powertrain/motor_controller"] + "reporting_sources": ["/powertrain/motor_controller"], + "owner": "/powertrain/motor_controller" } } diff --git a/postman/collections/ros2-medkit-gateway.postman_collection.json b/postman/collections/ros2-medkit-gateway.postman_collection.json index 5a138a0da..3fc2eefc1 100644 --- a/postman/collections/ros2-medkit-gateway.postman_collection.json +++ b/postman/collections/ros2-medkit-gateway.postman_collection.json @@ -1080,7 +1080,7 @@ "stream" ] }, - "description": "Server-Sent Events (SSE) stream for real-time fault notifications.\n\nEvents:\n- fault_confirmed: Fault transitioned to CONFIRMED status\n- fault_cleared: Fault ended (manually cleared, or auto-healed; check fault.status for CLEARED vs HEALED)\n- fault_updated: Fault data changed (occurrence_count, sources)\n\nSSE Format:\n```\nevent: fault_confirmed\ndata: {\"event_type\":\"fault_confirmed\",\"fault\":{...},\"timestamp\":1234567890.123}\n```\n\nNote: Postman has limited SSE support. For testing, use:\n- Browser: `new EventSource('http://localhost:8080/api/v1/faults/stream')`\n- curl: `curl -N {{base_url}}/faults/stream`\n\nFeatures:\n- Keepalive every 30 seconds\n- Reconnection support via Last-Event-ID header\n- Multiple simultaneous clients supported" + "description": "Server-Sent Events (SSE) stream for real-time fault notifications.\n\nEvents:\n- fault_confirmed: Fault transitioned to CONFIRMED status\n- fault_cleared: Fault ended (manually cleared, or auto-healed, check fault.status for CLEARED vs HEALED)\n- fault_updated: Record data changed (last_occurred, severity)\n\nSSE Format:\n```\nevent: fault_confirmed\ndata: {\"event_type\":\"fault_confirmed\",\"fault\":{...},\"timestamp\":1234567890.123}\n```\n\nNote: Postman has limited SSE support. For testing, use:\n- Browser: `new EventSource('http://localhost:8080/api/v1/faults/stream')`\n- curl: `curl -N {{base_url}}/faults/stream`\n\nFeatures:\n- Keepalive every 30 seconds\n- Reconnection support via Last-Event-ID header\n- Multiple simultaneous clients supported" }, "response": [] }, @@ -1142,7 +1142,7 @@ "faults" ] }, - "description": "List all faults for a component (REQ_INTEROP_012). Returns component_id, source_id (namespace), faults array, and count. By default returns PREFAILED and CONFIRMED faults." + "description": "List every fault record the component owns (REQ_INTEROP_012). A record is the pair (fault_code, reporting source), so two sources reporting one code are two items with distinct source_id. Returns component_id, faults array, and count. By default returns PREFAILED and CONFIRMED records." }, "response": [] }, @@ -1189,7 +1189,7 @@ "SENSOR_OVERTEMP" ] }, - "description": "Get a specific fault by fault_code (REQ_INTEROP_013). Returns component_id and fault object with fault_code, severity, severity_label, description, status, timestamps, and reporting_sources." + "description": "Get one fault record of this component by fault_code (REQ_INTEROP_013). Returns the SOVD detail whose x-medkit names the record's owner (x-medkit.owner) beside reporting_sources, occurrence_count and severity_label. Answers 409 (x-medkit-ambiguous-fault) when the code resolves to several records of this component, with parameters.owners naming them. Records the fault list shows rank above muted ones, and muted ones above cleared or healed ones, so one active record is served even while another source's cleared record of the code is kept." }, "response": [] }, @@ -1210,7 +1210,7 @@ "SENSOR_OVERTEMP" ] }, - "description": "Clear a fault by fault_code (REQ_INTEROP_015). Changes fault status to CLEARED. Returns success status with component_id and fault_code.\n\n**With correlation enabled:**\nIf the cleared fault is a root cause with `auto_clear_with_root=true`, the response includes `auto_cleared_codes[]` listing all symptom faults that were automatically cleared." + "description": "Clear one fault record of this component by fault_code (REQ_INTEROP_015). Changes that record's status to CLEARED and answers 204 with no body. Answers 409 (x-medkit-ambiguous-fault) when the code resolves to several records of this component, ranked the same way as the GET.\n\n**With correlation enabled:**\nThe per-entity route always sets skip_correlation_auto_clear, so symptom records correlated to the cleared one stay untouched." }, "response": [] }, diff --git a/src/ros2_medkit_fault_detection/README.md b/src/ros2_medkit_fault_detection/README.md index 55c9ec513..2265b9963 100644 --- a/src/ros2_medkit_fault_detection/README.md +++ b/src/ros2_medkit_fault_detection/README.md @@ -34,7 +34,10 @@ std::vector transitions = tracker.apply(signals); // raises + `evaluate` reports the full set of faults a rule governs (each flagged active or inactive). `FaultTransitionTracker` keeps the last-known state per `fault_code` and emits only the raise/clear edges, which a plugin forwards to -the fault manager as report/clear. +the fault manager as report/clear under the entity that owns the point. The +code-only key is the tracker's own: a fault manager record is identified by +`fault_code` and its reporting source, so one tracker per set of points sharing +a code, or unique codes within one tracker. ## Placement and packaging diff --git a/src/ros2_medkit_gateway/CMakeLists.txt b/src/ros2_medkit_gateway/CMakeLists.txt index ca40feb9a..937503f18 100644 --- a/src/ros2_medkit_gateway/CMakeLists.txt +++ b/src/ros2_medkit_gateway/CMakeLists.txt @@ -834,6 +834,9 @@ if(BUILD_TESTING) medkit_add_gtest(test_sse_fault_handler test/test_sse_fault_handler.cpp) target_link_libraries(test_sse_fault_handler gateway_ros2) medkit_target_dependencies(test_sse_fault_handler rclcpp ros2_medkit_msgs) + medkit_add_gtest(test_fault_handlers_plugin_clear test/test_fault_handlers_plugin_clear.cpp) + target_link_libraries(test_fault_handlers_plugin_clear gateway_ros2) + medkit_target_dependencies(test_fault_handlers_plugin_clear rclcpp ros2_medkit_msgs) # Zero-config entity freeze-frame capture tests medkit_add_gtest(test_entity_freeze_frame_capture test/test_entity_freeze_frame_capture.cpp) diff --git a/src/ros2_medkit_gateway/README.md b/src/ros2_medkit_gateway/README.md index 149c9c879..e91a14038 100644 --- a/src/ros2_medkit_gateway/README.md +++ b/src/ros2_medkit_gateway/README.md @@ -976,7 +976,15 @@ curl -X DELETE http://localhost:8080/api/v1/components/temp_sensor/configuration ### Faults Endpoints -Faults represent errors or warnings reported by system components. The gateway provides access to faults stored in `ros2_medkit_fault_manager`. +Faults represent errors or warnings reported by system components. The gateway provides access to fault records stored in `ros2_medkit_fault_manager`. + +One record is the pair (`fault_code`, reporting source). The source is the `source_id` the reporter used, it owns the record, and every fault item carries it as a top-level `source_id` beside the one-element `reporting_sources`. Two sources reporting one `fault_code` are two records, each with its own status, occurrence count and timestamps, and each cleared on its own. + +A per-entity route resolves a `fault_code` in the addressed entity's scope. The candidates come in three tiers, and the first tier holding any record of the code decides: the records the entity's fault list shows without a `status` parameter (`PREFAILED` and `CONFIRMED`, not muted), then the muted records of those statuses, then the records the list shows only for `status=cleared` or `status=healed` (`CLEARED`, `HEALED`, `PREPASSED`). So a muted or cleared record stays addressable by its code while nothing ranks above it, and once one of two sources has cleared its record the code names the other one. The per-code fault routes and the recording download by fault code all resolve this way. Exactly one candidate, and the route acts on it with its owner. Several, and it answers `409` with vendor error code `x-medkit-ambiguous-fault` and `parameters.owners` naming them, rather than acting on an arbitrary one. Address one of them through the route of the app that owns it, `/apps/{app_id}/faults/{fault_code}`. None, and it answers `404`. `DELETE /api/v1/{entity-path}/faults` clears the records the entity's fault list shows, each individually with its own owner. Like that list, it leaves muted records alone, and a muted record is cleared by its own per-code `DELETE`. + +A recording belongs to the fault records it is attached to, each a (`fault_code`, owner) pair. `GET /api/v1/{entity-path}/bulk-data/rosbags/{recording_id}` serves it when the owner of one of those records is in the entity's fault scope, the scope its fault list uses. Owning a record of the same code is not enough: two apps reporting one code each download only their own recording, cleared or not, and the component hosting both serves both. + +The detail response names the owner as `x-medkit.owner`. A fault list's `x-medkit.source_id` is a different thing, the addressed entity's own namespace path, so the two are never the same key. - `GET /api/v1/faults` - List all faults across the system (convenience API for dashboards) - `GET /api/v1/faults/stream` - Real-time fault event stream via Server-Sent Events (SSE) @@ -1038,8 +1046,8 @@ Real-time fault event stream using Server-Sent Events (SSE). Clients receive ins - **Real-time notifications**: Events pushed instantly when fault state changes - **Automatic reconnection**: Supports `Last-Event-ID` header for seamless reconnection - **Keepalive**: Sends `:keepalive` comment every 30 seconds to prevent timeouts -- **Event buffer**: Buffers up to 100 recent events for reconnecting clients. Under overflow the buffer evicts entries every live client has already received, then `fault_updated` entries superseded by a newer event for the same fault code - a lagging client still converges on the current state of every fault and loses no status transition. Anything beyond that (a superseded transition, or an event with no newer sibling) is genuinely lost for lagging clients: those losses are counted and logged as drops, and an affected client should refetch `GET /api/v1/faults` to resynchronize -- **Entity context (SOVD payload extension)**: When the gateway can resolve the fault's first reporting source back to an entity, the payload carries an `x-medkit` object with `entity_type` and `entity_id` fields so consumers can hit `/{entity_type}/{entity_id}/bulk-data/rosbags/{fault_code}` directly without enumerating entities. That address serves the fault's newest recording; to reach an older one, list `/bulk-data/rosbags` and use the descriptor `id` (the recording id) +- **Event buffer**: Buffers up to 100 recent events for reconnecting clients. Under overflow the buffer evicts entries every live client has already received, then `fault_updated` entries superseded by a newer event for the same fault record (the same code and owner) - a lagging client still converges on the current state of every fault and loses no status transition. Anything beyond that (a superseded transition, or an event with no newer sibling) is genuinely lost for lagging clients: those losses are counted and logged as drops, and an affected client should refetch `GET /api/v1/faults` to resynchronize +- **Entity context (SOVD payload extension)**: When the gateway can resolve the fault record's reporting source (its owner) to an entity, the payload carries an `x-medkit` object with `entity_type` and `entity_id` fields so consumers can hit `/{entity_type}/{entity_id}/bulk-data/rosbags/{fault_code}` directly without enumerating entities. That address serves the fault's newest recording; to reach an older one, list `/bulk-data/rosbags` and use the descriptor `id` (the recording id) - **Correlation payload**: when a root-cause event auto-clears correlated symptom faults, the payload carries their codes in `auto_cleared_codes` (omitted when empty); those symptoms get no event of their own **Event Types:** @@ -1083,7 +1091,7 @@ This replays any buffered events with ID > 5, then continues streaming new event #### GET /api/v1/components/{component_id}/faults -List all faults for a specific component. +List every fault record the component owns, through the apps it hosts plus, for an external component, its own id. **Query Parameters:** - `status` - Filter by fault status: `pending`, `confirmed`, `cleared`, `all` (default: `pending` + `confirmed`) diff --git a/src/ros2_medkit_gateway/config/examples/fault_owner_identity_manifest.yaml b/src/ros2_medkit_gateway/config/examples/fault_owner_identity_manifest.yaml new file mode 100644 index 000000000..bd01e2e16 --- /dev/null +++ b/src/ros2_medkit_gateway/config/examples/fault_owner_identity_manifest.yaml @@ -0,0 +1,53 @@ +# SOVD System Manifest: two owners of one fault code +# ================================================== +# Hybrid-mode manifest with TWO external apps hosted by one component. Both +# report the same fault_code to the fault_manager under their own entity ids, +# which makes two records of that code inside one component's fault scope. +# +# That shape is what the per-record identity has to answer for: each app's own +# route addresses its own record, the component's list shows both, and the +# component's per-code detail route has no single record to name and says so +# rather than picking one. +# +# Both apps are external: they report under their bare entity ids rather than a +# ROS FQN, which is also how a non-ROS asset bridged in by a protocol plugin +# reports. +# +# Used by: test/features/test_faults_owner_identity.test.py + +manifest_version: "1.0" + +metadata: + name: "fault-owner-identity" + version: "1.0.0" + description: "Two external apps under one component sharing a fault code" + +config: + unmanifested_nodes: warn + inherit_runtime_resources: true + +areas: + - id: shared-code-cell + name: "Shared Code Cell" + namespace: /shared_code_cell + description: "Cell hosting two assets that report the same fault code" + +components: + - id: shared-code-hub + name: "Shared Code Hub" + type: "controller" + area: shared-code-cell + description: "Hosts both reporting assets, so both records land in one scope" + +apps: + - id: owner-a + name: "Owner A" + external: true + is_located_on: shared-code-hub + description: "Reports SHARED_CODE under its own entity id" + + - id: owner-b + name: "Owner B" + external: true + is_located_on: shared-code-hub + description: "Reports the same SHARED_CODE under its own entity id" diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/faults/fault_scope.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/faults/fault_scope.hpp index f97102fbf..1908b8350 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/faults/fault_scope.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/faults/fault_scope.hpp @@ -17,6 +17,7 @@ #include #include #include +#include #include "ros2_medkit_gateway/core/models/entity_types.hpp" @@ -57,5 +58,49 @@ bool fault_in_source_scope(const nlohmann::json & fault, const std::set & source_fqns); +/// One fault record paired with the reporting source that owns it. +struct ScopedFault { + nlohmann::json fault; + std::string owner; +}; + +/// The reporting source that owns `fault`, or "" when the record names none. +/// A record carries exactly one, so this is its first (and only) entry. +std::string record_owner(const nlohmann::json & fault); + +/// Every record of `fault_code` in `faults_array` whose owner lies inside +/// `source_fqns`, ordered by owner so the answer does not depend on the order +/// the store listed them in. +/// +/// Returns the records rather than a count or a single pick, because the +/// consumers disagree about what several of them mean: a per-entity fault route +/// refuses to act on an ambiguous address, while a rosbag download asks each of +/// their owners whether it holds the recording. +std::vector records_of_code_in_scope(const nlohmann::json & faults_array, const std::string & fault_code, + const std::set & source_fqns); + +/// The records of `fault_code` inside `source_fqns` that a per-entity route +/// resolves a code in its URL over. +/// +/// `listing` is a fault-manager listing read with every status and with muted +/// records included: its `faults` array holds every record, and each +/// `muted_faults` entry names one the correlation engine hides by its +/// `fault_code` and `source_id`. +/// +/// The candidates come in three tiers, and the first tier holding any record of +/// the code decides: +/// 1. the records the entity's default fault list shows: PREFAILED and +/// CONFIRMED, not muted +/// 2. the muted records of those statuses, which that list would show but for +/// the correlation engine +/// 3. the records the list shows only when asked for status=cleared or +/// status=healed: CLEARED, HEALED and PREPASSED, muted or not +/// +/// So a record hidden as a symptom, or one a source cleared, stays addressable +/// by its code while it is the only one, and never makes the record a client +/// read off the list ambiguous. +std::vector addressable_records(const nlohmann::json & listing, const std::string & fault_code, + const std::set & source_fqns); + } // namespace faults } // namespace ros2_medkit_gateway diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/http/error_codes.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/http/error_codes.hpp index 988ee0220..358715156 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/http/error_codes.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/http/error_codes.hpp @@ -166,6 +166,12 @@ constexpr const char * ERR_SCRIPT_NOT_RUNNING = "x-medkit-script-not-running"; constexpr const char * ERR_SCRIPT_CONCURRENCY_LIMIT = "x-medkit-concurrency-limit"; constexpr const char * ERR_SCRIPT_FILE_TOO_LARGE = "x-medkit-script-too-large"; +/// A fault code addresses several records in the entity's scope, so the request +/// names no single record. Answered as 409 with the code and the owners in +/// `params`, because picking one arbitrarily is the failure the per-record +/// identity exists to remove. +constexpr const char * ERR_AMBIGUOUS_FAULT = "x-medkit-ambiguous-fault"; + /// Plugin provider returned an error (used for DataProvider/OperationProvider/FaultProvider errors) constexpr const char * ERR_PLUGIN_ERROR = "x-medkit-plugin-error"; diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/http/handlers/bulkdata_handlers.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/http/handlers/bulkdata_handlers.hpp index 5786d7222..941b8d5d4 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/http/handlers/bulkdata_handlers.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/http/handlers/bulkdata_handlers.hpp @@ -169,25 +169,42 @@ std::vector compute_bulkdata_source_filters(const ThreadSafeEntityC std::string rosbag_recording_id(const std::string & file_path); /** - * @brief Fault codes a rosbag download is authorized against. + * @brief Fault codes of the records a rosbag recording is attached to. * - * A recording is shared by every fault of a burst, so ownership is the union - * over those faults rather than a single code: the entity that owns any one of - * them may download the bag. That grants nothing new - before recordings had - * their own identity, each of those faults already addressed its own copy of - * the same bytes - it only renames the door. + * A recording is shared by every fault of a burst, so it belongs to several + * records, each a (fault code, owner) pair. These are the codes of those + * records. The download authorizes on the pairs, not on the codes: the codes + * only say which of the entity's records to ask about, and + * ``rosbag_rows_hold_recording`` says whether one of those records' owners + * really holds the recording. A code alone would let an entity owning any + * record of the code download another owner's recording of it. * * When the wire carries no ``fault_codes`` the response came from a peer that - * predates the field, where the addressed id *was* the fault code; authorizing - * against the requested id then reproduces the previous check exactly. + * predates the field, where the addressed id *was* the fault code, so the + * requested id stands in for it. * * @param rosbag_data Rosbag response from the fault manager * @param requested_id The ``{file_id}`` path segment the client asked for - * @return Non-empty list of fault codes to test against the entity's scope + * @return Non-empty list of fault codes whose in-scope records to check */ std::vector rosbag_attached_fault_codes(const nlohmann::json & rosbag_data, const std::string & requested_id); +/** + * @brief Does one owner's rosbag listing hold the recording that was served? + * + * ListRosbags answers the rows whose record is owned by exactly the source it + * was asked for, one row per (fault code, recording) link. A row naming the + * served recording therefore proves that a record of that owner is attached to + * it. A row names it by the same recording id, or, when either side predates + * recording ids, by the same bag path. + * + * @param rows The ``rosbags`` array of one owner's ListRosbags answer + * @param served Rosbag response from the fault manager for the download + * @return True when a row names the served recording + */ +bool rosbag_rows_hold_recording(const nlohmann::json & rows, const nlohmann::json & served); + /** * @brief Did the fault manager read the URL segment as a FAULT CODE rather than a * recording id? @@ -220,13 +237,24 @@ bool rosbag_resolved_by_fault_code(const nlohmann::json & rosbag_data, const std * Order follows first appearance, which is the order the fault manager listed * the rows in. * - * @param rows Rosbag rows as returned by the fault manager - * @param faults_by_code Faults keyed by code, for timestamp enrichment + * @param rows Rosbag rows as returned by the fault manager, each stamped with + * the ``owner`` it was listed under + * @param faults_by_record Fault records keyed by ``record_map_key``, for + * timestamp enrichment * @return One descriptor per distinct recording */ std::vector fold_rosbag_rows_into_descriptors(const std::vector & rows, - const std::unordered_map & faults_by_code); + const std::unordered_map & faults_by_record); + +/** + * @brief Key one fault record for the descriptor date lookup. + * + * A record is the pair (reporting source, fault code), and the two owners of + * one code have their own ``first_occurred``. Keying the lookup on the code + * alone let whichever record was listed last date the other's recordings. + */ +std::string record_map_key(const std::string & owner, const std::string & fault_code); } // namespace detail diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/managers/fault_manager.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/managers/fault_manager.hpp index 2b6627b7f..b7af783b9 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/managers/fault_manager.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/managers/fault_manager.hpp @@ -66,29 +66,37 @@ class FaultManager { const std::string & source_id); /// Get all faults, optionally filtered by component (prefix match on source_id). + /// The filter is applied gateway-side: ListFaults carries no source field. FaultResult list_faults(const std::string & source_id = "", bool include_prefailed = true, bool include_confirmed = true, bool include_cleared = false, bool include_healed = false, bool include_muted = false, bool include_clusters = false); - /// Get a specific fault by code with environment data, returned as JSON. + /// Get one fault record with environment data, returned as JSON. + /// A record is the pair (fault_code, owner), and `source_id` names the owner and + /// is sent in the request. Empty is unscoped and the fault manager refuses it + /// when several records carry the code. /// `data` carries `{ "fault": {...}, "environment_data": {...} }`. The /// rosbag-snapshot bulk_data_uri is intentionally NOT included; per-request /// URL building belongs to the handler that knows the entity path. FaultWithEnvJsonResult get_fault_with_env(const std::string & fault_code, const std::string & source_id = ""); - /// Get a specific fault by code (JSON result - "fault" body only). + /// Get one fault record (JSON result - "fault" body only). FaultResult get_fault(const std::string & fault_code, const std::string & source_id = ""); - /// Clear a fault. When `skip_correlation_auto_clear` is true the fault - /// manager will not cascade-clear correlated symptom fault codes - per-entity - /// DELETE routes set this to keep their clear inside the entity boundary. - FaultResult clear_fault(const std::string & fault_code, bool skip_correlation_auto_clear = false); + /// Clear one fault record, addressed by code and owning source. When + /// `skip_correlation_auto_clear` is true the fault manager will not + /// cascade-clear correlated symptom fault codes - per-entity DELETE routes + /// set this to keep their clear inside the entity boundary. + FaultResult clear_fault(const std::string & fault_code, const std::string & source_id, + bool skip_correlation_auto_clear = false); - /// Get snapshots for a fault (optional topic filter). - FaultResult get_snapshots(const std::string & fault_code, const std::string & topic = ""); + /// Get snapshots of one fault record (optional topic filter). + FaultResult get_snapshots(const std::string & fault_code, const std::string & source_id, + const std::string & topic = ""); - /// Get rosbag file info for a fault. - FaultResult get_rosbag(const std::string & fault_code); + /// Get rosbag file info for a recording id, or for one fault record when the + /// id is a fault code. `source_id` scopes the fault-code lookup only. + FaultResult get_rosbag(const std::string & id, const std::string & source_id = ""); /// Get all rosbag files for an entity (batch operation). FaultResult list_rosbags(const std::string & entity_fqn); diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/gateway_plugin.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/gateway_plugin.hpp index bb24aa58e..566ea6296 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/gateway_plugin.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/gateway_plugin.hpp @@ -122,16 +122,29 @@ class GatewayPlugin { } } + /// A copy of the log sink, for work that can outlive the call that started it. + /// + /// An asynchronous reply or a detached task that logs through `log_warn()` + /// holds `this`, and the plugin can be destroyed while that work runs. + /// Capturing this copy instead keeps the work from touching the plugin. The + /// copy is empty while no sink is wired, so check it before calling it. + std::function log_sink() const { + return log_fn_; + } + + /// Wire the log sink. PluginManager does this for every plugin it loads, + /// before configure(). A plugin hosted without a PluginManager, in a unit + /// test for example, wires its own through a subclass. The sink is not + /// synchronized with the log calls, so wire it while no plugin thread logs. + void set_logger(std::function fn) { + log_fn_ = std::move(fn); + } + private: friend class PluginManager; // Sets log_fn_ after construction /// Logging callback set by PluginManager. Routes to rclcpp::get_logger("plugin."). std::function log_fn_; - - /// Called by PluginManager to wire up logging - void set_logger(std::function fn) { - log_fn_ = std::move(fn); - } }; } // namespace ros2_medkit_gateway diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/plugin_manager.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/plugin_manager.hpp index 95b5534ac..d1fdeb378 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/plugin_manager.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/plugin_manager.hpp @@ -235,6 +235,29 @@ class PluginManager : public LogProviderRegistry { std::optional fetch_entity_data_via_route(const std::string & entity_id, const std::string & item = ""); + /** + * @brief The entity's current data content, from whichever source its owning + * plugin actually serves /data through. + * + * The owning plugin's `DataProvider::list_data` first, and only when the + * plugin exposes no provider for that entity, the plugin's own data route via + * `fetch_entity_data_via_route`. That is the order the /data endpoint itself + * resolves in, so a caller reading through this sees exactly what a client + * reading the entity would see. + * + * Reading the route first instead is not a fallback, it is a different + * answer: a plugin that serves /data through a provider and registers no + * vendor route returns nothing at all, and every caller downstream reads that + * as "this entity has no data right now". + * + * A throwing provider is treated as an absent one and falls through to the + * route: an in-process plugin call must not take the caller's loop down. + * + * @return Parsed content on success. Nullopt when the entity is not + * plugin-owned, or neither source can answer. + */ + std::optional fetch_entity_data_content(const std::string & entity_id); + /// Whether the entity's owning plugin registered a GET data route /// (x-plc-data) that fetch_entity_data_via_route() could dispatch. Cheap /// route-table check, no handler invocation - capability advertising uses diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/plugin_types.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/plugin_types.hpp index d4598fbb3..26c49efc3 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/plugin_types.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/plugins/plugin_types.hpp @@ -40,7 +40,18 @@ namespace ros2_medkit_gateway { /// so a pre-compiled v6 `.so` is rejected. Out-of-tree plugins must be /// recompiled against v7 headers; in-tree plugins that `return /// PLUGIN_API_VERSION` pick up the bump automatically. -constexpr int PLUGIN_API_VERSION = 7; +/// - v8: FaultProvider::clear_fault_record(entity_id, fault_code, owner), which +/// the gateway now calls to clear one fault record. Its default +/// implementation calls the unchanged two-argument clear_fault(), so +/// plugin SOURCE written against v7 compiles unchanged against v8 +/// headers (source-compatible). The new virtual changes the +/// FaultProvider vtable, so a pre-compiled v7 `.so` is rejected by the +/// strict equality check and must be recompiled against v8 headers. +/// GatewayPlugin also gains log_sink(), a copy of the log sink for work +/// that can outlive the plugin, and set_logger() becomes protected so a +/// plugin hosted without a PluginManager can wire its own. Both are +/// non-virtual and change no layout. +constexpr int PLUGIN_API_VERSION = 8; /// Log severity levels for plugin logging callback enum class PluginLogLevel { kInfo, kWarn, kError }; diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/providers/fault_provider.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/providers/fault_provider.hpp index 2f23d85f1..2f9b92aba 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/providers/fault_provider.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/providers/fault_provider.hpp @@ -89,6 +89,31 @@ class FaultProvider { /// @param fault_code Fault code to clear virtual tl::expected clear_fault(const std::string & entity_id, const std::string & fault_code) = 0; + + /// Clear one fault record, named by its code and the source that owns it. + /// + /// This is what the gateway calls for `DELETE /{entity}/faults/{code}` and + /// for each item of `DELETE /{entity}/faults`. A fault record is the pair + /// (fault_code, owner), so the code alone does not name one when several + /// sources report it. + /// + /// The default implementation calls `clear_fault(entity_id, fault_code)` and + /// drops the owner, so a provider written against the two-argument contract + /// compiles unchanged and keeps clearing by code. A provider that can scope + /// its clear to one record overrides this method. + /// + /// @param entity_id SOVD entity ID the request addressed + /// @param fault_code Fault code to clear + /// @param owner Reporting source that owns the record, as the gateway resolved + /// it in the entity's fault scope. It is NOT the entity id: a component + /// owns the records its hosted apps reported, so the owner is the app. + /// Empty only when the fault manager holds no record of this code (a + /// record the provider's own backend keeps, so the provider decides). + virtual tl::expected + clear_fault_record(const std::string & entity_id, const std::string & fault_code, const std::string & owner) { + static_cast(owner); + return clear_fault(entity_id, fault_code); + } }; } // namespace ros2_medkit_gateway diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/transports/fault_service_transport.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/transports/fault_service_transport.hpp index 8a4c1288c..53d1be498 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/transports/fault_service_transport.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/core/transports/fault_service_transport.hpp @@ -43,20 +43,30 @@ class FaultServiceTransport { bool include_cleared, bool include_healed, bool include_muted, bool include_clusters) = 0; + /// One fault record with its environment data. `source_id` names the source + /// that owns the record and is sent in the request, so the fault manager + /// resolves the record rather than the gateway filtering an arbitrary one. virtual FaultWithEnvJsonResult get_fault_with_env(const std::string & fault_code, const std::string & source_id) = 0; virtual FaultResult get_fault(const std::string & fault_code, const std::string & source_id) = 0; - /// Clear a fault by its fault_code. + /// Clear one fault record, addressed by `fault_code` and the `source_id` of + /// the source that owns it. An empty `source_id` is unscoped and the fault + /// manager refuses it when several records carry the code. /// `skip_correlation_auto_clear`, when true, asks the fault manager to NOT /// cascade-clear correlated symptom fault codes. Per-entity DELETE routes /// set this to true so the clear cannot reach faults reported by apps /// outside the addressed entity via the correlation graph. - virtual FaultResult clear_fault(const std::string & fault_code, bool skip_correlation_auto_clear = false) = 0; + virtual FaultResult clear_fault(const std::string & fault_code, const std::string & source_id, + bool skip_correlation_auto_clear = false) = 0; - virtual FaultResult get_snapshots(const std::string & fault_code, const std::string & topic) = 0; + virtual FaultResult get_snapshots(const std::string & fault_code, const std::string & source_id, + const std::string & topic) = 0; - virtual FaultResult get_rosbag(const std::string & fault_code) = 0; + /// Rosbag info for one recording, or for one fault record when `id` is a + /// fault code. `source_id` names the record's owner and scopes the fault-code + /// lookup only. The recording-id path ignores it. + virtual FaultResult get_rosbag(const std::string & id, const std::string & source_id) = 0; virtual FaultResult list_rosbags(const std::string & entity_fqn) = 0; diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/dto/faults.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/dto/faults.hpp index 64e07a49d..67ef3aa3c 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/dto/faults.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/dto/faults.hpp @@ -43,7 +43,7 @@ namespace dto { // Wire keys (exact, from fault_msg_conversions.cpp): // fault_code, severity, description, first_occurred, last_occurred, // last_passed (absent = never passed), occurrence_count, status, -// reporting_sources, severity_label +// reporting_sources, source_id, severity_label // ============================================================================= struct FaultListItem { std::string fault_code; @@ -55,6 +55,10 @@ struct FaultListItem { std::optional occurrence_count; std::string status; std::optional> reporting_sources; + /// The reporting source that owns this record. Together with `fault_code` it + /// addresses the record on the per-record routes, and it is the single entry + /// of `reporting_sources`. + std::optional source_id; std::optional severity_label; // enum: INFO|WARN|ERROR|CRITICAL|UNKNOWN }; @@ -64,7 +68,7 @@ inline constexpr auto dto_fields = std::make_tuple( field("description", &FaultListItem::description), field("first_occurred", &FaultListItem::first_occurred), field("last_occurred", &FaultListItem::last_occurred), field("last_passed", &FaultListItem::last_passed), field("occurrence_count", &FaultListItem::occurrence_count), field("status", &FaultListItem::status), - field("reporting_sources", &FaultListItem::reporting_sources), + field("reporting_sources", &FaultListItem::reporting_sources), field("source_id", &FaultListItem::source_id), field_enum("severity_label", &FaultListItem::severity_label, kFaultSeverityLabelValues)); template <> @@ -242,11 +246,20 @@ inline constexpr std::string_view dto_name = "FaultListAggX // FaultXMedkit - x-medkit vendor extension inside FaultDetail // // Wire keys (from build_sovd_fault_response): -// occurrence_count, reporting_sources, severity_label, status_raw +// occurrence_count, reporting_sources, owner, severity_label, status_raw +// +// The key is `owner`, not `source_id`, deliberately. A fault LIST's x-medkit +// already carries `source_id` meaning the addressed entity's namespace path +// (FaultListXMedkit below), so reusing that name here for the record's +// reporting source would give one key two meanings inside one API. // ============================================================================= struct FaultXMedkit { std::optional occurrence_count; std::optional> reporting_sources; + /// The reporting source that owns this record. It is what the per-record + /// routes take as their `source_id` request field, and the single entry of + /// `reporting_sources`. + std::optional owner; std::optional severity_label; std::optional status_raw; }; @@ -254,7 +267,7 @@ struct FaultXMedkit { template <> inline constexpr auto dto_fields = std::make_tuple(field("occurrence_count", &FaultXMedkit::occurrence_count), - field("reporting_sources", &FaultXMedkit::reporting_sources), + field("reporting_sources", &FaultXMedkit::reporting_sources), field("owner", &FaultXMedkit::owner), field("severity_label", &FaultXMedkit::severity_label), field("status_raw", &FaultXMedkit::status_raw)); @@ -513,11 +526,10 @@ struct SchemaWriter { {"additionalProperties", true}, {"x-medkit-opaque", true}, {"description", - "Acknowledgement of a clear, shaped by whoever owns the entity. The ROS 2 path answers " - "`{\"code\": , \"cleared\": true}`; a plugin answers with its backend's own " - "acknowledgement - UDS clear response codes, vendor warnings, residual fault state - and the " - "gateway emits it verbatim. Treat the 2xx status, not a body field, as the signal that the clear " - "succeeded."}}; + "Acknowledgement of a clear, shaped by whoever owns the entity. The ROS 2 path answers 204 with " + "no body at all. A plugin answers with its backend's own acknowledgement (UDS clear response " + "codes, vendor warnings, residual fault state) and the gateway emits it verbatim. Treat the 2xx " + "status, not a body field, as the signal that the clear succeeded."}}; } }; diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/dto/sse_frames.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/dto/sse_frames.hpp index 2d827783d..93319f84b 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/dto/sse_frames.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/dto/sse_frames.hpp @@ -98,9 +98,9 @@ inline constexpr std::string_view dto_name = "TriggerEventFra // FaultStreamXMedkit - the `x-medkit` object on a fault-stream frame. // // Present only when the fault's reporting source resolves to a known entity. -// It is a hint for addressing the fault's bulk-data, not an ownership claim: -// a debounced fault can have several co-reporters and this names the -// lexicographically first. +// An event describes one record, and a record has exactly one reporting +// source, its owner, so this names the entity that owner resolves to: the one +// whose routes address that record and its bulk-data. // ----------------------------------------------------------------------------- struct FaultStreamXMedkit { std::string entity_type; diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/entity_freeze_frame_capture.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/entity_freeze_frame_capture.hpp index 7dc081dc0..f77699058 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/entity_freeze_frame_capture.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/entity_freeze_frame_capture.hpp @@ -123,8 +123,9 @@ class EntityFreezeFrameCapture { EntityFreezeFrameCapture(EntityFreezeFrameCapture &&) = delete; EntityFreezeFrameCapture & operator=(EntityFreezeFrameCapture &&) = delete; - /// Frames captured for a fault code (empty when none). Thread-safe. - std::vector frames_for(const std::string & fault_code) const; + /// Frames captured for one fault record, addressed by its code and the + /// reporting source that owns it (empty when none). Thread-safe. + std::vector frames_for(const std::string & fault_code, const std::string & owner) const; /// Build the compact {resource_id: value} dict from a DataProvider::list_data /// response. Items without a "value" field map to null; a response without an @@ -208,10 +209,12 @@ class EntityFreezeFrameCapture { /// Guards frames_, insertion_order_ and fallback_logged_: capture thread /// writes, HTTP handler threads read. mutable std::mutex mutex_; - /// Keyed by fault_code only: cross-entity isolation relies on the - /// fault_manager keeping reporting_sources append-only for a code and on - /// get_fault gating by source scope. A per-source clear upstream would need - /// per-entity eviction here too. + /// Keyed by the record the frames belong to, `(fault_code, owner)`. A fault + /// code alone is not a record: two sources reporting one code confirm + /// independently, and under a code-only key the second confirmation + /// overwrote the first source's frames and served them on its detail page. + /// The key is rendered as `fault_code + '\0' + owner`, which no id can + /// contain, so the two halves cannot run together. std::unordered_map> frames_; std::deque insertion_order_; ///< eviction order (FIFO) std::unordered_set fallback_logged_; ///< fault codes already warned about (bounded) diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/http/handlers/fault_handlers.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/http/handlers/fault_handlers.hpp index e7667f229..bd21c38f8 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/http/handlers/fault_handlers.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/http/handlers/fault_handlers.hpp @@ -17,10 +17,12 @@ #include #include #include +#include #include #include #include +#include "ros2_medkit_gateway/core/faults/fault_scope.hpp" #include "ros2_medkit_gateway/core/faults/fault_types.hpp" #include "ros2_medkit_gateway/dto/faults.hpp" #include "ros2_medkit_gateway/entity_freeze_frame_capture.hpp" @@ -159,22 +161,18 @@ class FaultHandlers { * or is a strict path-child (i.e. `/<...>`), so similarly named nodes * like `/ns/node` and `/ns/node_extra` are not conflated. * - * The "all sources must match" semantic (rather than "any source") is - * deliberate: it blocks two cross-entity escalation paths. - * - * 1. GET would otherwise return a response whose `reporting_sources` and - * environment data carry identities of nodes the caller has no - * business reading. - * 2. DELETE would otherwise let a viewer of entity A clear the aggregated - * fault record for a fault that entity B also reports, because the - * underlying `ClearFault.srv` has no scope argument. + * A record names one reporting source, its owner, so in practice this asks + * whether that owner is in scope. The "all sources must match" form is kept + * because the gateway also serves records relayed from peers and read back + * from older stores, and a record that somehow carries a source outside the + * entity must stay invisible to it rather than half-visible. * * An empty scope set, an empty `reporting_sources` array, a missing * `reporting_sources` field, or any non-string source entry all return * false - there is no vacuous "all match" case. * - * Public for direct unit testing; called by `get_fault`, `clear_fault`, - * and indirectly via the per-entity collection routes. + * Public for direct unit testing. Called by `resolve_scoped_fault` and + * indirectly via the per-entity collection routes. */ static bool fault_in_source_scope(const nlohmann::json & fault, const std::set & source_fqns); @@ -219,7 +217,56 @@ class FaultHandlers { static nlohmann::json merge_entity_freeze_frames(nlohmann::json env_data, const std::vector & frames); + /** + * @brief Map every reporting source the cache can attribute to the entity + * that owns it, for the global clear's lock check. + * + * A source is whatever a reporter put in `source_id`, which for an external + * app or an external component is its bare SOVD id and not a ROS FQN. The map + * was built from `App::effective_fqn()` alone, which is empty for exactly + * those entities, so `DELETE /faults` found none of them and cleared their + * records straight through another client's lock - while the locking document + * said the route skips a locked entity's faults. + * + * Public static for direct unit testing. Called by `clear_all_faults_global`. + */ + static std::unordered_map build_source_entity_map(const ThreadSafeEntityCache & cache); + + /** + * @brief Turn the candidate records of one code into the single record a + * per-record route may act on, or the error it answers instead. + * + * `records` are the candidates `faults::addressable_records` yields: the + * records of the code the entity's fault list shows, else the muted ones, + * else the cleared or healed ones. A fault code addresses as many records as + * there are sources reporting it, so two candidates mean the entity has not + * named one. The route answers 409 `x-medkit-ambiguous-fault` with those owners in + * `parameters.owners` rather than acting on whichever record the store + * happened to list first: picking one is the failure the per-record identity + * exists to remove, and the caller cannot tell from the response which one it + * got. No candidate is the 404 the routes have always answered. + * + * Public static for direct unit testing. Called by `resolve_scoped_fault` and + * by the plugin clear. + */ + static tl::expected select_scoped_fault(std::vector records, + const std::string & fault_code, + const std::string & id_field, + const std::string & entity_id); + private: + /** + * @brief Resolve the one record of `fault_code` that `entity` owns. + * + * Lists every status with muted records included, narrows them to this + * code's candidates with `faults::addressable_records` (the records the + * entity's list shows, else its muted ones, else its cleared or healed ones) + * and hands those to `select_scoped_fault`. The owner it returns is what the + * per-record services are then called with. + */ + tl::expected resolve_scoped_fault(const EntityInfo & entity_info, + const std::string & fault_code); + HandlerContext & ctx_; }; diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/http/handlers/sse_fault_handler.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/http/handlers/sse_fault_handler.hpp index 79c0f4d2d..732eb19bd 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/http/handlers/sse_fault_handler.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/http/handlers/sse_fault_handler.hpp @@ -68,10 +68,11 @@ namespace handlers { * - Replay buffer of up to 100 events. Eviction order under overflow: * 1. entries every live client has already been sent (free), * 2. fault_updated entries superseded by a newer event for the same fault - * code (coalesced - a lagging client still converges on the correct - * current state, and no status transition is erased), + * record, the same code and owner (coalesced - a lagging client still + * converges on the correct current state, and no status transition is + * erased), * 3. transition entries (confirmed / cleared) superseded by a newer event - * for the same code - the current state survives but the transition + * for the same record - the current state survives but the transition * history is lost, so this IS counted and logged as a drop, * 4. the oldest entry, owed and not superseded - counted and logged too. */ @@ -153,9 +154,9 @@ class SSEFaultHandler { /** * @brief Events genuinely lost: evicted while still owed to a live client, - * except fault_updated entries a newer same-code event supersedes. Includes - * superseded transitions - their history is gone even though the current - * state survives. + * except fault_updated entries a newer event of the same record supersedes. + * Includes superseded transitions - their history is gone even though the + * current state survives. * * Rotation of the replay buffer with no client attached is NOT a drop. */ @@ -256,7 +257,8 @@ class SSEFaultHandler { /// live client is attached. Caller holds queue_mutex_. std::optional delivered_watermark_locked() const; - /// Oldest buffered entry whose state a later same-code entry supersedes. + /// Oldest buffered entry whose state a later entry of the same record (fault + /// code and owner, from the same peer) supersedes. /// With `updates_only` set, only fault_updated entries qualify (erasing them /// cannot erase a status transition). Entries carrying auto_cleared_codes /// are never chosen: that correlation payload exists nowhere else. Caller diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/ros2/conversions/fault_msg_conversions.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/ros2/conversions/fault_msg_conversions.hpp index 2869f394f..e816f2038 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/ros2/conversions/fault_msg_conversions.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/ros2/conversions/fault_msg_conversions.hpp @@ -25,7 +25,8 @@ namespace ros2_medkit_gateway::ros2::conversions { /// /// Produces the flat per-fault representation consumed by `list_faults` items, /// the SSE fault-event payload, and the trigger subsystem's notifier change -/// values. Timestamps become seconds-as-double; severity gains a label. +/// values. Timestamps become seconds-as-double, severity gains a label, and +/// `source_id` names the reporting source that owns the record. /// /// Lives at the ROS-coupled boundary because three independent call sites /// each translate `ros2_medkit_msgs::msg::Fault` directly into JSON: the diff --git a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/ros2/transports/ros2_fault_service_transport.hpp b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/ros2/transports/ros2_fault_service_transport.hpp index 6c408ca5e..37f5ef2bf 100644 --- a/src/ros2_medkit_gateway/include/ros2_medkit_gateway/ros2/transports/ros2_fault_service_transport.hpp +++ b/src/ros2_medkit_gateway/include/ros2_medkit_gateway/ros2/transports/ros2_fault_service_transport.hpp @@ -74,11 +74,13 @@ class Ros2FaultServiceTransport : public FaultServiceTransport { FaultResult get_fault(const std::string & fault_code, const std::string & source_id) override; - FaultResult clear_fault(const std::string & fault_code, bool skip_correlation_auto_clear) override; + FaultResult clear_fault(const std::string & fault_code, const std::string & source_id, + bool skip_correlation_auto_clear) override; - FaultResult get_snapshots(const std::string & fault_code, const std::string & topic) override; + FaultResult get_snapshots(const std::string & fault_code, const std::string & source_id, + const std::string & topic) override; - FaultResult get_rosbag(const std::string & fault_code) override; + FaultResult get_rosbag(const std::string & id, const std::string & source_id) override; FaultResult list_rosbags(const std::string & entity_fqn) override; diff --git a/src/ros2_medkit_gateway/src/core/fault_trigger_engine.cpp b/src/ros2_medkit_gateway/src/core/fault_trigger_engine.cpp index 2e376a27b..ef6c4e74e 100644 --- a/src/ros2_medkit_gateway/src/core/fault_trigger_engine.cpp +++ b/src/ros2_medkit_gateway/src/core/fault_trigger_engine.cpp @@ -176,16 +176,20 @@ tl::expected> FaultTriggerEngine:: } std::lock_guard lock(mutex_); - // fault_code is the PRIMARY KEY of the fault-manager store, so two rules - // sharing one code would fight over the same stored fault (one rule's clear - // erases the other's assertion) and make per-app delete/clear ambiguous. + // One rule per fault_code across every app, checked with no app filter. The + // fault manager keys a record by (fault_code, reporting source), so two rules + // on DIFFERENT apps sharing a code would own two separate records and not + // collide there. The reason the engine still refuses it is its own: a code is + // how an operator names a rule's condition, and one code asserted by two + // rules cannot be read back to either of them. const auto dup = std::find_if(rules_.begin(), rules_.end(), [&](const FaultTriggerRule & r) { return r.fault_code == rule.fault_code; }); if (dup != rules_.end()) { - return tl::make_unexpected( - std::make_pair(409, "fault_code '" + rule.fault_code + "' is already used by rule '" + dup->id + "' on app '" + - dup->app_id + "' - fault codes are global to the fault store, pick a distinct one")); + return tl::make_unexpected(std::make_pair( + 409, "fault_code '" + rule.fault_code + "' is already used by rule '" + dup->id + "' on app '" + dup->app_id + + "'. The trigger engine keeps one rule per fault code across every app, so " + "that a code names exactly one rule's condition. Pick a distinct fault_code.")); } rule.id = "ftr_" + std::to_string(next_seq_++); rules_.push_back(rule); diff --git a/src/ros2_medkit_gateway/src/core/faults/fault_scope.cpp b/src/ros2_medkit_gateway/src/core/faults/fault_scope.cpp index 0f377026d..1646d9be5 100644 --- a/src/ros2_medkit_gateway/src/core/faults/fault_scope.cpp +++ b/src/ros2_medkit_gateway/src/core/faults/fault_scope.cpp @@ -14,6 +14,7 @@ #include "ros2_medkit_gateway/core/faults/fault_scope.hpp" +#include #include #include @@ -91,6 +92,16 @@ void collect_function_app_fqns(const ThreadSafeEntityCache & cache, const std::s } } +/// True when the entity's fault list shows `fault` when no `status` is asked +/// for. That list reads the fault manager with parse_fault_status_param's +/// defaults, pending and confirmed, which the transport sends as PREFAILED and +/// CONFIRMED. CLEARED, HEALED and PREPASSED appear only under status=cleared, +/// status=healed or status=all. +bool shown_by_default_list(const nlohmann::json & fault) { + const auto status = fault.value("status", std::string{}); + return status == "PREFAILED" || status == "CONFIRMED"; +} + bool source_matches_scope(const std::string & src, const std::set & scope_fqns) { for (const auto & fqn : scope_fqns) { if (src == fqn) { @@ -180,5 +191,91 @@ nlohmann::json filter_faults_by_sources(const nlohmann::json & faults_array, return filtered; } +std::string record_owner(const nlohmann::json & fault) { + if (!fault.contains("reporting_sources") || !fault["reporting_sources"].is_array()) { + return {}; + } + const auto & sources = fault["reporting_sources"]; + if (sources.empty() || !sources.front().is_string()) { + return {}; + } + return sources.front().get(); +} + +std::vector records_of_code_in_scope(const nlohmann::json & faults_array, const std::string & fault_code, + const std::set & source_fqns) { + std::vector records; + if (!faults_array.is_array()) { + return records; + } + for (const auto & fault : faults_array) { + if (!fault.is_object() || fault.value("fault_code", std::string{}) != fault_code) { + continue; + } + // The scope predicate the list routes filter with. It decides which entity + // a record belongs to, not whether that entity's list shows it: the list + // also leaves muted, cleared and healed records out, which is why a + // per-record route resolves a code through addressable_records (the records + // the list shows first, then muted ones, then cleared or healed ones) + // rather than over this result as it stands. + if (!fault_in_source_scope(fault, source_fqns)) { + continue; + } + records.push_back(ScopedFault{fault, record_owner(fault)}); + } + std::sort(records.begin(), records.end(), [](const ScopedFault & a, const ScopedFault & b) { + return a.owner < b.owner; + }); + return records; +} + +std::vector addressable_records(const nlohmann::json & listing, const std::string & fault_code, + const std::set & source_fqns) { + if (!listing.is_object()) { + return {}; + } + auto records = records_of_code_in_scope(listing.value("faults", nlohmann::json::array()), fault_code, source_fqns); + + // Owners whose record of this code is muted. An entry is one muted record, + // so the owner has to match as well as the code: muting one owner's record + // never hides another owner's record of the same code. + std::set muted_owners; + const auto muted_it = listing.find("muted_faults"); + if (muted_it != listing.end() && muted_it->is_array()) { + for (const auto & entry : *muted_it) { + if (entry.is_object() && entry.value("fault_code", std::string{}) == fault_code) { + muted_owners.insert(entry.value("source_id", std::string{})); + } + } + } + + // Three tiers, and the first one holding any record decides. The records the + // entity's default fault list shows come first. Then the muted ones, which + // that list would show but for the correlation engine. Then everything the + // list hides for its status: CLEARED, HEALED and PREPASSED, muted or not. + // Without the last split a record one source cleared long ago stayed a + // candidate beside the record the list shows, and the code answered 409 for + // as long as the cleared record was kept. + std::vector shown; + std::vector muted; + std::vector inactive; + for (auto & record : records) { + if (!shown_by_default_list(record.fault)) { + inactive.push_back(std::move(record)); + } else if (muted_owners.count(record.owner) > 0) { + muted.push_back(std::move(record)); + } else { + shown.push_back(std::move(record)); + } + } + if (!shown.empty()) { + return shown; + } + if (!muted.empty()) { + return muted; + } + return inactive; +} + } // namespace faults } // namespace ros2_medkit_gateway diff --git a/src/ros2_medkit_gateway/src/core/managers/fault_manager.cpp b/src/ros2_medkit_gateway/src/core/managers/fault_manager.cpp index cb02ef227..bc40550d5 100644 --- a/src/ros2_medkit_gateway/src/core/managers/fault_manager.cpp +++ b/src/ros2_medkit_gateway/src/core/managers/fault_manager.cpp @@ -41,16 +41,18 @@ FaultResult FaultManager::get_fault(const std::string & fault_code, const std::s return transport_->get_fault(fault_code, source_id); } -FaultResult FaultManager::clear_fault(const std::string & fault_code, bool skip_correlation_auto_clear) { - return transport_->clear_fault(fault_code, skip_correlation_auto_clear); +FaultResult FaultManager::clear_fault(const std::string & fault_code, const std::string & source_id, + bool skip_correlation_auto_clear) { + return transport_->clear_fault(fault_code, source_id, skip_correlation_auto_clear); } -FaultResult FaultManager::get_snapshots(const std::string & fault_code, const std::string & topic) { - return transport_->get_snapshots(fault_code, topic); +FaultResult FaultManager::get_snapshots(const std::string & fault_code, const std::string & source_id, + const std::string & topic) { + return transport_->get_snapshots(fault_code, source_id, topic); } -FaultResult FaultManager::get_rosbag(const std::string & fault_code) { - return transport_->get_rosbag(fault_code); +FaultResult FaultManager::get_rosbag(const std::string & id, const std::string & source_id) { + return transport_->get_rosbag(id, source_id); } FaultResult FaultManager::list_rosbags(const std::string & entity_fqn) { diff --git a/src/ros2_medkit_gateway/src/entity_freeze_frame_capture.cpp b/src/ros2_medkit_gateway/src/entity_freeze_frame_capture.cpp index fe060826c..115b5a3ad 100644 --- a/src/ros2_medkit_gateway/src/entity_freeze_frame_capture.cpp +++ b/src/ros2_medkit_gateway/src/entity_freeze_frame_capture.cpp @@ -97,10 +97,20 @@ EntityFreezeFrameCapture::~EntityFreezeFrameCapture() { subscription_slot_.reset(); } -std::vector -EntityFreezeFrameCapture::frames_for(const std::string & fault_code) const { +namespace { + +/// The storage key for one record. '\0' cannot appear in a fault code or in a +/// reporting source, so no pair of ids can collide with another pair. +std::string record_key(const std::string & fault_code, const std::string & owner) { + return fault_code + '\0' + owner; +} + +} // namespace + +std::vector EntityFreezeFrameCapture::frames_for(const std::string & fault_code, + const std::string & owner) const { std::lock_guard lock(mutex_); - auto it = frames_.find(fault_code); + auto it = frames_.find(record_key(fault_code, owner)); return it != frames_.end() ? it->second : std::vector{}; } @@ -314,16 +324,21 @@ void EntityFreezeFrameCapture::capture_standing_faults() { RCLCPP_WARN(logger_, "Entity freeze-frame startup catch-up failed: standing-fault lister threw"); return; } - // Codes with a confirm already queued belong to the drain loop: capturing - // them here too would read the plugin twice for one confirm. - std::unordered_set queued_codes; + // Records with a confirm already queued belong to the drain loop: capturing + // them here too would read the plugin twice for one confirm. The key is the + // record and not the code, because another source's record of the same code + // has its own confirm and its own frame, and skipping it on the code alone + // would leave that record with no frame at all. + std::unordered_set queued_records; { std::lock_guard lock(queue_mutex_); if (stop_) { return; } for (const auto & queued : queue_) { - queued_codes.insert(queued->fault.fault_code); + for (const auto & source : queued->fault.reporting_sources) { + queued_records.insert(record_key(queued->fault.fault_code, source)); + } } } size_t framed = 0; @@ -335,8 +350,14 @@ void EntityFreezeFrameCapture::capture_standing_faults() { if (fault.fault_code.empty() || fault.reporting_sources.empty()) { continue; } - if (queued_codes.count(fault.fault_code) != 0) { - continue; + std::vector pending_sources; + for (const auto & source : fault.reporting_sources) { + if (queued_records.count(record_key(fault.fault_code, source)) == 0) { + pending_sources.push_back(source); + } + } + if (pending_sources.empty()) { + continue; // every record of this code is already on the drain loop } if (framed >= max_faults_) { ++over_cap; // storing more would FIFO-evict this catch-up's own frames @@ -345,7 +366,7 @@ void EntityFreezeFrameCapture::capture_standing_faults() { ros2_medkit_msgs::msg::FaultEvent event; event.event_type = ros2_medkit_msgs::msg::FaultEvent::EVENT_CONFIRMED; event.fault.fault_code = fault.fault_code; - event.fault.reporting_sources = fault.reporting_sources; + event.fault.reporting_sources = std::move(pending_sources); if (capture_for_event(event, /*startup_catchup=*/true)) { ++framed; } @@ -386,6 +407,10 @@ bool EntityFreezeFrameCapture::capture_for_event(const ros2_medkit_msgs::msg::Fa bool startup_catchup) { const std::string & fault_code = event.fault.fault_code; + // The event describes one record, so its reporting_sources names one owner. + // The loop is kept because a record read back from an older store, or relayed + // from a peer that predates the per-record identity, can still carry several, + // and every one of them is framed under the key it owns. std::vector frames; for (const auto & source : event.fault.reporting_sources) { DataProvider * provider = resolver_ ? resolver_(source) : nullptr; @@ -427,15 +452,23 @@ bool EntityFreezeFrameCapture::capture_for_event(const ros2_medkit_msgs::msg::Fa } } + // One entry per record, under the record's owner. An event describes one + // record, so `frames` holds one frame and its entity_id IS that owner. A + // record read back from an older store or relayed from a peer can still carry + // several sources, and each of those is stored under the source that reported + // it rather than merged, because frames_for asks for one owner at a time. std::lock_guard lock(mutex_); - if (frames_.find(fault_code) == frames_.end()) { - insertion_order_.push_back(fault_code); - while (frames_.size() >= max_faults_ && !insertion_order_.empty()) { - frames_.erase(insertion_order_.front()); - insertion_order_.pop_front(); + for (auto & frame : frames) { + const auto key = record_key(fault_code, frame.entity_id); + if (frames_.find(key) == frames_.end()) { + insertion_order_.push_back(key); + while (frames_.size() >= max_faults_ && !insertion_order_.empty()) { + frames_.erase(insertion_order_.front()); + insertion_order_.pop_front(); + } } + frames_[key] = std::vector{frame}; } - frames_[fault_code] = std::move(frames); RCLCPP_DEBUG(logger_, "Captured entity freeze-frame(s) for fault '%s'", fault_code.c_str()); return true; diff --git a/src/ros2_medkit_gateway/src/gateway_node.cpp b/src/ros2_medkit_gateway/src/gateway_node.cpp index b17256644..23a4e2544 100644 --- a/src/ros2_medkit_gateway/src/gateway_node.cpp +++ b/src/ros2_medkit_gateway/src/gateway_node.cpp @@ -1976,13 +1976,21 @@ void GatewayNode::init_fault_trigger_engine() { } const int poll_ms = static_cast(poll_ms_used); - // Value fetch: same in-process x-plc-data route the freeze-frame capture uses, - // so a rule always sees exactly the value the /data endpoint would serve. - auto fetcher = [this](const std::string & app_id, const std::string & data_name) -> std::optional { - if (!plugin_mgr_) { - return std::nullopt; - } - const auto content = plugin_mgr_->fetch_entity_data_via_route(app_id); + // The owning plugin's own DataProvider first, its vendor route only as a + // fallback. This is the order the freeze-frame capture and trigger creation + // already use, and reading the route first meant a rule on an app whose + // plugin serves /data through a provider and exposes no vendor route could + // never evaluate: the fetch returned nullopt every tick and the rule sat + // there holding its state, silently doing nothing. + auto entity_data_content = [this](const std::string & app_id) -> std::optional { + return plugin_mgr_ ? plugin_mgr_->fetch_entity_data_content(app_id) : std::nullopt; + }; + + // Value fetch: the same content the /data endpoint would serve, so a rule + // always sees exactly the value a client reading that app would see. + auto fetcher = [entity_data_content](const std::string & app_id, + const std::string & data_name) -> std::optional { + const auto content = entity_data_content(app_id); // A down link serves frozen last-known values; nullopt holds rule state // (fault_trigger_engine) instead of firing on a stale number all outage. if (!content || !EntityFreezeFrameCapture::content_has_live_data(*content) || @@ -2014,24 +2022,26 @@ void GatewayNode::init_fault_trigger_engine() { fault_mgr_->report_fault(fault_code, sev, description, app_id); }; - auto clear = [this](const std::string & /*app_id*/, const std::string & fault_code) { + auto clear = [this](const std::string & app_id, const std::string & fault_code) { if (fault_mgr_) { - // A trigger-rule auto-clear (falling edge / rule deletion) is a local, - // rule-scoped event: skip the correlation engine's cascade so clearing - // one rule's fault can never wipe correlated faults owned by anything - // else. fault_code is globally unique (create() rejects duplicates), so - // the app scope is already encoded in the code itself. - fault_mgr_->clear_fault(fault_code, /*skip_correlation_auto_clear=*/true); + // The record this rule raised is the one it reported under: the report + // lambda above passes app_id as source_id, so the clear names the same + // owner. Clearing by code alone would reach another source's record of + // the same code, which a rule on this app has no business touching. + // + // A trigger-rule auto-clear (falling edge / rule deletion) is also a + // local, rule-scoped event: skip the correlation engine's cascade so + // clearing one rule's fault can never wipe correlated faults owned by + // anything else. + fault_mgr_->clear_fault(fault_code, app_id, /*skip_correlation_auto_clear=*/true); } }; - // Data-point enumeration for create-time validation: same in-process route as - // the fetcher, so "exists" means exactly "the fetcher could ever read it". - auto data_point_names = [this](const std::string & app_id) -> std::optional> { - if (!plugin_mgr_) { - return std::nullopt; - } - const auto content = plugin_mgr_->fetch_entity_data_via_route(app_id); + // Data-point enumeration for create-time validation: the same content source + // as the fetcher, in the same order, so "exists" means exactly "the fetcher + // could ever read it". + auto data_point_names = [entity_data_content](const std::string & app_id) -> std::optional> { + const auto content = entity_data_content(app_id); if (!content || !EntityFreezeFrameCapture::content_has_live_data(*content)) { return std::nullopt; } diff --git a/src/ros2_medkit_gateway/src/http/handlers/bulkdata_handlers.cpp b/src/ros2_medkit_gateway/src/http/handlers/bulkdata_handlers.cpp index df98c93f8..afbd4e0e3 100644 --- a/src/ros2_medkit_gateway/src/http/handlers/bulkdata_handlers.cpp +++ b/src/ros2_medkit_gateway/src/http/handlers/bulkdata_handlers.cpp @@ -143,6 +143,31 @@ std::vector rosbag_attached_fault_codes(const nlohmann::json & rosb return {requested_id}; } +bool rosbag_rows_hold_recording(const nlohmann::json & rows, const nlohmann::json & served) { + if (!rows.is_array() || !served.is_object()) { + return false; + } + const std::string served_id = served.value("recording_id", std::string{}); + const std::string served_path = served.value("file_path", std::string{}); + for (const auto & row : rows) { + if (!row.is_object()) { + continue; + } + const std::string row_id = row.value("recording_id", std::string{}); + if (!served_id.empty() && !row_id.empty()) { + if (row_id == served_id) { + return true; + } + continue; + } + // One side predates recording ids. The bag path is the recording then. + if (!served_path.empty() && row.value("file_path", std::string{}) == served_path) { + return true; + } + } + return false; +} + bool rosbag_resolved_by_fault_code(const nlohmann::json & rosbag_data, const std::string & requested_id) { const std::string resolved = rosbag_data.value("recording_id", ""); // An absent id means a peer that predates the field; it answers by fault code @@ -150,9 +175,14 @@ bool rosbag_resolved_by_fault_code(const nlohmann::json & rosbag_data, const std return !resolved.empty() && resolved != requested_id; } +std::string record_map_key(const std::string & owner, const std::string & fault_code) { + // '\0' cannot appear in either id, so the two halves cannot run together. + return owner + '\0' + fault_code; +} + std::vector fold_rosbag_rows_into_descriptors(const std::vector & rows, - const std::unordered_map & faults_by_code) { + const std::unordered_map & faults_by_record) { struct RecordingEntry { std::string recording_id; std::string format; @@ -184,7 +214,8 @@ fold_rosbag_rows_into_descriptors(const std::vector & rows, // fault only for a row that predates the field. int64_t created_at_ns = row.value("created_at_ns", int64_t{0}); if (created_at_ns == 0) { - if (auto it = faults_by_code.find(fault_code); it != faults_by_code.end()) { + const auto key = record_map_key(row.value("owner", std::string{}), fault_code); + if (auto it = faults_by_record.find(key); it != faults_by_record.end()) { const double first_occurred = it->second.value("first_occurred", 0.0); created_at_ns = static_cast(first_occurred * 1'000'000'000); } @@ -335,27 +366,37 @@ BulkDataHandlers::list_descriptors(const http::TypedRequest & req) { // Functions aggregate rosbags from all hosting apps. auto source_filters = get_source_filters(entity); - // Collect faults across all source filters for timestamp enrichment + // Records for the date fallback, keyed by the record they belong to rather + // than by code alone: two owners of one code are two records with their own + // first_occurred, and keying on the code let whichever was listed last date + // the other's recordings. Every status, because an acknowledged record still + // has its recordings served. std::unordered_map fault_map; - for (const auto & source_filter : source_filters) { - auto faults_result = fault_mgr->list_faults(source_filter); + { + auto faults_result = fault_mgr->list_faults("", /*include_prefailed=*/true, /*include_confirmed=*/true, + /*include_cleared=*/true, /*include_healed=*/true, + /*include_muted=*/false, /*include_clusters=*/false); if (faults_result.success && faults_result.data.contains("faults")) { for (const auto & fault_json : faults_result.data["faults"]) { if (fault_json.contains("fault_code")) { - std::string fc = fault_json["fault_code"].get(); - fault_map[fc] = fault_json; + fault_map[detail::record_map_key(faults::record_owner(fault_json), + fault_json["fault_code"].get())] = fault_json; } } } } - // Collect rosbags across all source filters + // Collect rosbags across all source filters. ListRosbags matches the record + // owner exactly, so every row it answers with is owned by the filter it was + // asked under. The row itself does not carry the owner, so stamp it here + // while that is still known. std::vector all_rosbags; for (const auto & source_filter : source_filters) { auto rosbags_result = fault_mgr->list_rosbags(source_filter); if (rosbags_result.success && rosbags_result.data.contains("rosbags")) { - for (const auto & rosbag : rosbags_result.data["rosbags"]) { - all_rosbags.push_back(rosbag); + for (auto rosbag : rosbags_result.data["rosbags"]) { + rosbag["owner"] = source_filter; + all_rosbags.push_back(std::move(rosbag)); } } } @@ -440,43 +481,98 @@ http::Result BulkDataHandlers::download(const http::TypedR // is exactly one authorization semantic, not two. auto fault_mgr = ctx_.node()->get_fault_manager(); - auto rosbag_result = fault_mgr->get_rosbag(bulk_data_id); + auto source_filters = get_source_filters(entity); + std::set scope(source_filters.begin(), source_filters.end()); + + // Every record this entity owns, in one read, every status and muted ones + // too. Both the code resolution and the authorization below ask which + // RECORDS a code addresses, not whether a code exists: two sources reporting + // one code are two records, and an unscoped read of that code is ambiguous + // and answers nothing at all. A muted or cleared record keeps its + // recordings, so leaving it out would 404 them. Reading the list once also + // replaces one GetFault round trip per attached code. + auto held = fault_mgr->list_faults("", /*include_prefailed=*/true, /*include_confirmed=*/true, + /*include_cleared=*/true, /*include_healed=*/true, /*include_muted=*/true, + /*include_clusters=*/false); + const json listing = held.success ? held.data : json::object(); + const json all_faults = listing.value("faults", json::array()); + + // A compatibility URL carries a fault code, so the entity's own records of + // that code say which owner to ask the recording for. The same rule as the + // fault routes picks them: the records the entity's fault list shows, then + // its muted ones, then its cleared or healed ones, the first of those that + // holds any. An id that is really a recording id matches none and the owner + // stays empty, which is what the recording path ignores anyway. + // + // Several candidates means the URL names none of them, and each owner keeps + // its own recordings. Serving the lowest-sorting owner's bag would hand the + // caller another owner's bytes under a URL that never said whose they were, + // so this refuses the same way the fault routes do. The recording-id form + // is unaffected: it addresses the bag directly and needs no owner. + auto requested_records = faults::addressable_records(listing, bulk_data_id, scope); + if (requested_records.size() > 1) { + std::vector owners; + owners.reserve(requested_records.size()); + for (const auto & record : requested_records) { + owners.push_back(record.owner); + } + return tl::unexpected( + make_error(409, ERR_AMBIGUOUS_FAULT, "Fault code addresses several records in this entity", + json{{"details", + "Several sources this entity owns report this fault code, and each record keeps its own " + "recordings. parameters.owners names them. Address the recording by its own id: the " + "rosbags listing of this entity (GET .../bulk-data/rosbags) carries it as each " + "descriptor's id, and a fault detail links it as " + "environment_data.snapshots[].bulk_data_uri."}, + {"entity_id", path_info->entity_id}, + {"fault_code", bulk_data_id}, + {"owners", owners}})); + } + const std::string requested_owner = requested_records.empty() ? std::string{} : requested_records.front().owner; + + auto rosbag_result = fault_mgr->get_rosbag(bulk_data_id, requested_owner); if (!rosbag_result.success || !rosbag_result.data.contains("file_path")) { return tl::unexpected( make_error(404, ERR_RESOURCE_NOT_FOUND, "Bulk-data not found", json{{"bulk_data_id", bulk_data_id}})); } - // Security check: the bag belongs to this entity when ANY fault it was captured - // for is within the entity's source scope. Union rather than a single code - // because a burst shares one recording, and each of those faults already had its - // own downloadable copy of it before - so this grants nothing new, it only - // renames the door. Tested with the shared boundary-aware matcher: the - // transport's get_fault(code, source) check is a raw prefix match, so app id - // "plc" would otherwise claim the assets of "plc_line1". - auto source_filters = get_source_filters(entity); - std::set scope(source_filters.begin(), source_filters.end()); - // The compatibility path needs the REQUESTED code in scope, not just some code // the recording is attached to. A burst shares one bag, so authorizing on the // union alone answers 200 for a fault code this entity does not own - the bytes // are ones it could already fetch under its own code, but the 200 itself tells - // the caller that another fault shares its recording. build_sovd_fault_response - // refuses to mix sources for the same reason. When the id really was a recording - // id there is no requested code and the union is the whole answer. - if (detail::rosbag_resolved_by_fault_code(rosbag_result.data, bulk_data_id)) { - auto requested = fault_mgr->get_fault(bulk_data_id, ""); - if (!requested.success || !faults::fault_in_source_scope(requested.data, scope)) { - return tl::unexpected(make_error(404, ERR_RESOURCE_NOT_FOUND, "Bulk-data not found for this entity", - json{{"entity_id", path_info->entity_id}})); - } + // the caller that another fault shares its recording. + if (detail::rosbag_resolved_by_fault_code(rosbag_result.data, bulk_data_id) && requested_records.empty()) { + return tl::unexpected(make_error(404, ERR_RESOURCE_NOT_FOUND, "Bulk-data not found for this entity", + json{{"entity_id", path_info->entity_id}})); } const auto attached_codes = detail::rosbag_attached_fault_codes(rosbag_result.data, bulk_data_id); - const bool authorized = std::any_of(attached_codes.begin(), attached_codes.end(), [&](const std::string & code) { - auto fault_result = fault_mgr->get_fault(code, ""); - return fault_result.success && faults::fault_in_source_scope(fault_result.data, scope); - }); + // The bag belongs to this entity when one of the records it is attached to, + // a (fault code, owner) pair, has its owner in scope. The code alone does not + // say that: two sources reporting one code are two records with their own + // recordings, and owning one of them, cleared or not, does not make the + // other's recording this entity's. GetRosbag names the attached codes but + // not their owners, so the candidates are the owners of this entity's own + // records of those codes, and ListRosbags, which answers exactly the rows + // one owner holds, says whether the served recording is among them: a + // recording is downloadable exactly when a listing under one of those owners + // carries it. The owners are matched with the fault list's scope rule, so + // an owner below one of the entity's sources counts too. + std::set candidate_owners; + for (const auto & code : attached_codes) { + for (const auto & record : faults::records_of_code_in_scope(all_faults, code, scope)) { + if (!record.owner.empty()) { + candidate_owners.insert(record.owner); + } + } + } + const bool authorized = + std::any_of(candidate_owners.begin(), candidate_owners.end(), [&](const std::string & owner) { + auto held_rows = fault_mgr->list_rosbags(owner); + return held_rows.success && + detail::rosbag_rows_hold_recording(held_rows.data.value("rosbags", json::array()), rosbag_result.data); + }); if (!authorized) { return tl::unexpected(make_error(404, ERR_RESOURCE_NOT_FOUND, "Bulk-data not found for this entity", json{{"entity_id", path_info->entity_id}})); diff --git a/src/ros2_medkit_gateway/src/http/handlers/fault_handlers.cpp b/src/ros2_medkit_gateway/src/http/handlers/fault_handlers.cpp index b235acc53..6821e5cf1 100644 --- a/src/ros2_medkit_gateway/src/http/handlers/fault_handlers.cpp +++ b/src/ros2_medkit_gateway/src/http/handlers/fault_handlers.cpp @@ -228,6 +228,88 @@ ErrorInfo FaultHandlers::classify_fault_failure(FaultFailure failure, const std: return make_error(503, ERR_SERVICE_UNAVAILABLE, unavailable_summary, params); } +std::unordered_map +FaultHandlers::build_source_entity_map(const ThreadSafeEntityCache & cache) { + std::unordered_map source_to_entity; + for (const auto & app : cache.get_apps()) { + // The same resolution the fault scope uses: an external app's bare id, any + // other app's effective FQN, and nothing for an unbound non-external app. + auto source = faults::resolve_app_source_fqn(cache, app.id); + if (!source.empty()) { + source_to_entity[source] = app.id; + } + } + for (const auto & component : cache.get_components()) { + // Only an external component claims its bare id as a reporting source. + if (component.external.value_or(false) && !component.id.empty()) { + source_to_entity.emplace(component.id, component.id); + } + } + return source_to_entity; +} + +tl::expected +FaultHandlers::select_scoped_fault(std::vector records, const std::string & fault_code, + const std::string & id_field, const std::string & entity_id) { + if (records.empty()) { + return tl::make_unexpected( + make_error(404, ERR_RESOURCE_NOT_FOUND, "Fault not found", + json{{"details", + "No fault record with this code is reported by a source this entity owns. A record is the " + "pair (fault_code, reporting source), so a code another entity's source reports is not this " + "entity's record."}, + {id_field, entity_id}, + {"fault_code", fault_code}})); + } + if (records.size() > 1) { + std::vector owners; + owners.reserve(records.size()); + for (const auto & record : records) { + owners.push_back(record.owner); + } + return tl::make_unexpected( + make_error(409, ERR_AMBIGUOUS_FAULT, "Fault code addresses several records in this entity", + json{{"details", + "Several sources this entity owns report this fault code, and each is its own record. " + "parameters.owners names them. Address one through the route of the app that owns it, " + "/apps/{app_id}/faults/{fault_code}."}, + {id_field, entity_id}, + {"fault_code", fault_code}, + {"owners", owners}})); + } + return std::move(records.front()); +} + +tl::expected FaultHandlers::resolve_scoped_fault(const EntityInfo & entity_info, + const std::string & fault_code) { + auto * fault_mgr = ctx_.node()->get_fault_manager(); + if (fault_mgr == nullptr) { + return tl::make_unexpected(make_error(503, ERR_SERVICE_UNAVAILABLE, "Failed to get fault", + json{{entity_info.id_field, entity_info.id}, {"fault_code", fault_code}})); + } + + // Every status, and muted records too. A record the caller addresses by code + // exists whatever its lifecycle state, and the detail and clear routes have + // always served a cleared or healed one. Muting is the correlation engine + // hiding a symptom from the entity's list, not a reason the record stops + // being addressable. addressable_records resolves over the records that list + // shows first, then the muted ones, then the cleared or healed ones, so + // neither a muted record nor one a source cleared turns the record a client + // read off the list into an ambiguous address. + auto result = fault_mgr->list_faults("", /*include_prefailed=*/true, /*include_confirmed=*/true, + /*include_cleared=*/true, /*include_healed=*/true, /*include_muted=*/true, + /*include_clusters=*/false); + if (!result.success) { + return tl::make_unexpected(classify_fault_failure(result.failure, result.error_message, "Failed to get fault", + entity_info.id_field, entity_info.id, fault_code)); + } + + const auto & cache = ctx_.node()->get_thread_safe_cache(); + auto source_fqns = HandlerContext::resolve_entity_source_fqns(cache, entity_info); + auto records = faults::addressable_records(result.data, fault_code, source_fqns); + return select_scoped_fault(std::move(records), fault_code, entity_info.id_field, entity_info.id); +} + bool FaultHandlers::fault_in_source_scope(const json & fault, const std::set & source_fqns) { // Thin wrapper preserving the public static API; the scope logic now lives in // the neutral core helper shared with the ROS 2 plugin-context fault path. @@ -420,6 +502,11 @@ dto::FaultDetail FaultHandlers::build_sovd_fault_response(const json & fault_jso dto::FaultXMedkit xm; xm.occurrence_count = static_cast(fault_json.value("occurrence_count", static_cast(0))); if (!reporting_sources.empty()) { + // The owner, beside the list it is the single entry of: a client reading + // the detail can address the record it is looking at without unpacking an + // array. Named `owner` and not `source_id` because a fault LIST's x-medkit + // already uses `source_id` for the addressed entity's namespace path. + xm.owner = reporting_sources.front(); xm.reporting_sources = std::move(reporting_sources); } xm.severity_label = severity_to_label(severity); @@ -761,16 +848,21 @@ http::Result FaultHandlers::get_fault(const http::TypedR // freeze-frame/rosbag snapshots, consistent with the non-plugin detail // path. Fall through to the plugin's own provider for faults the // fault_manager does not hold (e.g. on-demand UDS DTCs). - if (auto * fault_mgr = ctx_.node()->get_fault_manager(); fault_mgr != nullptr) { - auto mgr_result = fault_mgr->get_fault_with_env(fault_code, ""); - if (mgr_result.success) { - const auto & owned_fault_json = mgr_result.data.value("fault", json::object()); - const auto & cache = ctx_.node()->get_thread_safe_cache(); - auto source_fqns = HandlerContext::resolve_entity_source_fqns(cache, entity_info); - if (FaultHandlers::fault_in_source_scope(owned_fault_json, source_fqns)) { + auto scoped = resolve_scoped_fault(entity_info, fault_code); + if (!scoped && scoped.error().http_status == 409) { + // Several of this entity's own sources report the code. That is an + // answer, not a miss, so it must not fall through to the plugin. + return tl::make_unexpected(scoped.error()); + } + if (scoped) { + if (auto * fault_mgr = ctx_.node()->get_fault_manager(); fault_mgr != nullptr) { + auto mgr_result = fault_mgr->get_fault_with_env(fault_code, scoped->owner); + if (mgr_result.success) { + const auto & owned_fault_json = mgr_result.data.value("fault", json::object()); json env_data_json = mgr_result.data.value("environment_data", json::object()); if (auto * capture = ctx_.node()->get_entity_freeze_frame_capture()) { - env_data_json = merge_entity_freeze_frames(std::move(env_data_json), capture->frames_for(fault_code)); + env_data_json = merge_entity_freeze_frames(std::move(env_data_json), + capture->frames_for(fault_code, scoped->owner)); } auto detail = build_sovd_fault_response(owned_fault_json, env_data_json, entity_path_info->entity_path); return wrap_detail_result(dto::JsonWriter::write(detail)); @@ -798,9 +890,15 @@ http::Result FaultHandlers::get_fault(const http::TypedR } } - auto fault_mgr = ctx_.node()->get_fault_manager(); + // Which record. The entity's scope decides, and the owner it yields is what + // the enriched read is then addressed with. + auto scoped = resolve_scoped_fault(entity_info, fault_code); + if (!scoped) { + return tl::make_unexpected(scoped.error()); + } - auto result = fault_mgr->get_fault_with_env(fault_code, ""); + auto fault_mgr = ctx_.node()->get_fault_manager(); + auto result = fault_mgr->get_fault_with_env(fault_code, scoped->owner); if (!result.success) { return tl::make_unexpected(classify_fault_failure(result.failure, result.error_message, "Failed to get fault", entity_info.id_field, entity_id, fault_code)); @@ -809,22 +907,10 @@ http::Result FaultHandlers::get_fault(const http::TypedR // Build SOVD-compliant response from the transport-supplied JSON shape. const auto & fault_json = result.data.value("fault", json::object()); - const auto & cache = ctx_.node()->get_thread_safe_cache(); - auto source_fqns = HandlerContext::resolve_entity_source_fqns(cache, entity_info); - if (!FaultHandlers::fault_in_source_scope(fault_json, source_fqns)) { - return tl::make_unexpected( - make_error(404, ERR_RESOURCE_NOT_FOUND, "Fault not found", - json{{"details", - "Fault is not in scope for this entity: every reporting source must be one of the entity's " - "owned apps, and a mixed-source fault that includes any out-of-entity reporter is rejected " - "to prevent cross-entity disclosure"}, - {entity_info.id_field, entity_id}, - {"fault_code", fault_code}})); - } - json env_data_json = result.data.value("environment_data", json::object()); if (auto * capture = ctx_.node()->get_entity_freeze_frame_capture()) { - env_data_json = merge_entity_freeze_frames(std::move(env_data_json), capture->frames_for(fault_code)); + env_data_json = + merge_entity_freeze_frames(std::move(env_data_json), capture->frames_for(fault_code, scoped->owner)); } auto detail = build_sovd_fault_response(fault_json, env_data_json, entity_path_info->entity_path); @@ -888,35 +974,51 @@ FaultHandlers::clear_fault(const http::TypedRequest & req) { // native path): fall through to the scope-checked fault_manager clear // below like any other entity. if (fault_prov != nullptr) { - // Cross-entity clear guard: the plugin clear_fault forwards the code to - // the fault_manager with no scope check, so without this a - // DELETE /{A}/faults/{code} could clear a fault owned by entity B. When - // the fault_manager holds this fault, require it to be in this entity's - // source scope before delegating; reject out-of-scope. Faults the - // fault_manager does not hold (plugin-internal, e.g. on-demand UDS DTCs) - // fall through to the plugin provider unchanged. Mirrors the ownership - // check in the plugin get_fault branch and the non-plugin clear path. - if (auto * fault_mgr = ctx_.node()->get_fault_manager(); fault_mgr != nullptr) { - auto mgr_result = fault_mgr->get_fault_with_env(fault_code, ""); - if (mgr_result.success) { - const auto & owned_fault_json = mgr_result.data.value("fault", json::object()); - const auto & cache = ctx_.node()->get_thread_safe_cache(); - auto source_fqns = HandlerContext::resolve_entity_source_fqns(cache, entity_info); - if (!FaultHandlers::fault_in_source_scope(owned_fault_json, source_fqns)) { - return tl::make_unexpected( - make_error(404, ERR_RESOURCE_NOT_FOUND, "Fault not found", - json{{"details", - "Fault is not in scope for this entity: every reporting source must be one of the " - "entity's owned apps, and a mixed-source fault that includes any out-of-entity " - "reporter is rejected to prevent cross-entity clear"}, - {entity_info.id_field, entity_id}, - {"fault_code", fault_code}})); - } + // Cross-entity clear guard: the plugin clear_fault names its own entity + // but the record it reaches is still addressed by code, so without this + // a DELETE /{A}/faults/{code} could reach a record owned by entity B. + // When the fault_manager holds a record of this code at all, require one + // in this entity's scope before delegating. Records the fault_manager + // does not hold (plugin-internal, e.g. on-demand UDS DTCs) fall through + // to the plugin provider unchanged. + // + // The owner the gateway resolved travels to the provider, because that + // is which record this route addresses. It stays empty only when the + // fault manager answered and holds no record of this code at all, the + // plugin-internal case the provider decides for itself. A fault manager + // that cannot be read is not that case: the route cannot tell a record + // it holds from one it does not, so it answers 503 as the native path + // does and never calls the provider with an owner it did not resolve. + auto * fault_mgr = ctx_.node()->get_fault_manager(); + if (fault_mgr == nullptr) { + return tl::make_unexpected(make_error(503, ERR_SERVICE_UNAVAILABLE, "Failed to clear fault", + json{{entity_info.id_field, entity_id}, {"fault_code", fault_code}})); + } + auto held = fault_mgr->list_faults("", /*include_prefailed=*/true, /*include_confirmed=*/true, + /*include_cleared=*/true, /*include_healed=*/true, + /*include_muted=*/true, /*include_clusters=*/false); + if (!held.success) { + return tl::make_unexpected(classify_fault_failure(held.failure, held.error_message, "Failed to clear fault", + entity_info.id_field, entity_id, fault_code)); + } + std::string resolved_owner; + const auto & all = held.data.value("faults", json::array()); + const bool store_holds_code = std::any_of(all.begin(), all.end(), [&](const json & fault) { + return fault.is_object() && fault.value("fault_code", std::string{}) == fault_code; + }); + if (store_holds_code) { + const auto & cache = ctx_.node()->get_thread_safe_cache(); + auto source_fqns = HandlerContext::resolve_entity_source_fqns(cache, entity_info); + auto scoped = select_scoped_fault(faults::addressable_records(held.data, fault_code, source_fqns), fault_code, + entity_info.id_field, entity_id); + if (!scoped) { + return tl::make_unexpected(scoped.error()); } + resolved_owner = scoped->owner; } try { - auto result = fault_prov->clear_fault(entity_id, fault_code); + auto result = fault_prov->clear_fault_record(entity_id, fault_code, resolved_owner); if (!result) { return tl::make_unexpected( make_plugin_error(result.error().http_status, result.error().message, json{{"entity_id", entity_id}})); @@ -937,29 +1039,13 @@ FaultHandlers::clear_fault(const http::TypedRequest & req) { auto fault_mgr = ctx_.node()->get_fault_manager(); - // Verify the fault is in this entity's scope BEFORE clearing. - auto get_result = fault_mgr->get_fault_with_env(fault_code, ""); - if (!get_result.success) { - return tl::make_unexpected(classify_fault_failure(get_result.failure, get_result.error_message, - "Failed to clear fault", entity_info.id_field, entity_id, - fault_code)); + // Which record this entity means, before anything is cleared. + auto scoped = resolve_scoped_fault(entity_info, fault_code); + if (!scoped) { + return tl::make_unexpected(scoped.error()); } - const auto & cache = ctx_.node()->get_thread_safe_cache(); - auto source_fqns = HandlerContext::resolve_entity_source_fqns(cache, entity_info); - const auto & fault_json = get_result.data.value("fault", json::object()); - if (!FaultHandlers::fault_in_source_scope(fault_json, source_fqns)) { - return tl::make_unexpected( - make_error(404, ERR_RESOURCE_NOT_FOUND, "Fault not found", - json{{"details", - "Fault is not in scope for this entity: every reporting source must be one of the entity's " - "owned apps, and a mixed-source fault that includes any out-of-entity reporter is rejected " - "to prevent cross-entity disclosure"}, - {entity_info.id_field, entity_id}, - {"fault_code", fault_code}})); - } - - auto result = fault_mgr->clear_fault(fault_code, /*skip_correlation_auto_clear=*/true); + auto result = fault_mgr->clear_fault(fault_code, scoped->owner, /*skip_correlation_auto_clear=*/true); if (!result.success) { return tl::make_unexpected(classify_fault_failure(result.failure, result.error_message, "Failed to clear fault", entity_info.id_field, entity_id, fault_code)); @@ -1017,7 +1103,14 @@ http::Result FaultHandlers::clear_all_faults(const http::TypedR if (code.empty()) { continue; } - auto clear_result = fault_prov->clear_fault(entity_id, code); + // Each listed item names its own record. Sending the entity id + // for all of them addresses at most one owner's record and + // silently leaves the others standing. + auto owner = fault.value("source_id", std::string{}); + if (owner.empty()) { + owner = faults::record_owner(fault); + } + auto clear_result = fault_prov->clear_fault_record(entity_id, code, owner); if (!clear_result) { failed_codes.push_back(code); } @@ -1048,6 +1141,9 @@ http::Result FaultHandlers::clear_all_faults(const http::TypedR // resolves through `HandlerContext::resolve_entity_source_fqns` so the // area BFS, function-hosting-component expansion, and wildcard-app // empty-set behavior stay consistent across all four fault routes. + // The defaults leave muted records out, as the entity's fault list does. + // That is deliberate: this route clears what the list shows, and a muted + // record is cleared by its own per-code DELETE, which resolves it. auto result = fault_mgr->list_faults(""); if (!result.success) { return tl::make_unexpected( @@ -1058,8 +1154,10 @@ http::Result FaultHandlers::clear_all_faults(const http::TypedR auto entity_fqns = HandlerContext::resolve_entity_source_fqns(cache, entity_info); json faults_to_clear = faults::filter_faults_by_sources(result.data["faults"], entity_fqns); - // Clear each matching fault. Use `skip_correlation_auto_clear=true` for - // the same reason as the single-fault DELETE: keep this entity's clear + // Clear each in-scope RECORD, each with its own owner: two sources of this + // entity reporting one code are two records, and clearing by code alone + // would leave one of them standing. Use `skip_correlation_auto_clear=true` + // for the same reason as the single-fault DELETE: keep this entity's clear // from cascading into correlated symptoms reported by other entities. if (faults_to_clear.is_array()) { for (const auto & fault : faults_to_clear) { @@ -1067,10 +1165,11 @@ http::Result FaultHandlers::clear_all_faults(const http::TypedR continue; } std::string code = fault["fault_code"].get(); - auto clear_result = fault_mgr->clear_fault(code, /*skip_correlation_auto_clear=*/true); + std::string owner = faults::record_owner(fault); + auto clear_result = fault_mgr->clear_fault(code, owner, /*skip_correlation_auto_clear=*/true); if (!clear_result.success) { - RCLCPP_WARN(HandlerContext::logger(), "Failed to clear fault '%s' for entity '%s': %s", code.c_str(), - entity_id.c_str(), clear_result.error_message.c_str()); + RCLCPP_WARN(HandlerContext::logger(), "Failed to clear fault '%s' of source '%s' for entity '%s': %s", + code.c_str(), owner.c_str(), entity_id.c_str(), clear_result.error_message.c_str()); } } } @@ -1108,41 +1207,35 @@ FaultHandlers::clear_all_faults_global(const http::TypedRequest & req) { json{{"details", faults_result.error_message}})); } - // Build FQN-to-entity-ID map for lock checking + // Build source-to-entity-ID map for lock checking. A source is whatever a + // reporter put in `source_id`, which for an external app or an external + // component is its bare SOVD id and not a ROS FQN. Keying on + // `effective_fqn()` alone therefore found none of them, and the lock of + // every external entity - every protocol bridge, every PLC - went + // unhonoured on this route while the document said it was honoured. auto * lock_mgr = ctx_.node() ? ctx_.node()->get_lock_manager() : nullptr; - std::unordered_map fqn_to_entity; + std::unordered_map source_to_entity; if (lock_mgr) { - const auto & cache = ctx_.node()->get_thread_safe_cache(); - for (const auto & app : cache.get_apps()) { - auto fqn = app.effective_fqn(); - if (!fqn.empty()) { - fqn_to_entity[fqn] = app.id; - } - } + source_to_entity = build_source_entity_map(ctx_.node()->get_thread_safe_cache()); } auto client_id = req.header("X-Client-Id").value_or(std::string{}); - // Clear each fault, skipping those on locked entities + // Clear each RECORD, each with its owner, skipping records whose owning + // entity is locked by another client. if (faults_result.data.contains("faults") && faults_result.data["faults"].is_array()) { for (const auto & fault : faults_result.data["faults"]) { if (!fault.contains("fault_code")) { continue; } - // Check if any reporting source is on a locked entity + const std::string owner = faults::record_owner(fault); bool blocked = false; - if (lock_mgr && fault.contains("reporting_sources")) { - for (const auto & src : fault["reporting_sources"]) { - auto src_str = src.get(); - auto it = fqn_to_entity.find(src_str); - if (it != fqn_to_entity.end()) { - auto access = lock_mgr->check_access(it->second, client_id, "faults"); - if (!access.allowed) { - blocked = true; - break; - } - } + if (lock_mgr) { + auto it = source_to_entity.find(owner); + if (it != source_to_entity.end()) { + auto access = lock_mgr->check_access(it->second, client_id, "faults"); + blocked = !access.allowed; } } @@ -1151,10 +1244,10 @@ FaultHandlers::clear_all_faults_global(const http::TypedRequest & req) { } std::string code = fault["fault_code"].get(); - auto clear_result = fault_mgr->clear_fault(code); + auto clear_result = fault_mgr->clear_fault(code, owner); if (!clear_result.success) { - RCLCPP_WARN(HandlerContext::logger(), "Failed to clear fault '%s': %s", code.c_str(), - clear_result.error_message.c_str()); + RCLCPP_WARN(HandlerContext::logger(), "Failed to clear fault '%s' of source '%s': %s", code.c_str(), + owner.c_str(), clear_result.error_message.c_str()); } } } diff --git a/src/ros2_medkit_gateway/src/http/handlers/sse_fault_handler.cpp b/src/ros2_medkit_gateway/src/http/handlers/sse_fault_handler.cpp index 00d4816d4..81d401ee5 100644 --- a/src/ros2_medkit_gateway/src/http/handlers/sse_fault_handler.cpp +++ b/src/ros2_medkit_gateway/src/http/handlers/sse_fault_handler.cpp @@ -149,12 +149,16 @@ std::optional SSEFaultHandler::delivered_watermark_locked() const { } std::deque::iterator SSEFaultHandler::find_superseded_locked(bool updates_only) { - // Keyed on the peer as well as the code: a fault code is unique on the - // gateway that raised it and nowhere else, so two peers reporting the same - // code are reporting two faults, and treating one as the newer state of the - // other would delete a live fault from a client's view. + // Keyed on the record, not the code: a record is (fault_code, owner), so two + // sources reporting one code are two faults, and one owner's newer event says + // nothing about the other owner's record. Keyed on the peer as well: a fault + // is unique on the gateway that raised it and nowhere else, so two peers + // reporting the same record are reporting two faults. Treating either as the + // newer state of the other would delete a live fault from a client's view. auto supersede_key = [](const QueuedEvent & queued) { - return queued.peer + '\0' + queued.event.fault.fault_code; + const auto & sources = queued.event.fault.reporting_sources; + const std::string owner = sources.empty() ? std::string{} : sources.front(); + return queued.peer + '\0' + queued.event.fault.fault_code + '\0' + owner; }; std::unordered_map newest_index; for (std::size_t i = 0; i < event_queue_.size(); ++i) { @@ -215,8 +219,8 @@ SSEFaultHandler::EvictionStats SSEFaultHandler::evict_to_capacity_locked() { ++stats.coalesced; continue; } - // Something a live client is owed has to go. Prefer an entry a newer - // same-code event supersedes: the current state still reaches the client + // Something a live client is owed has to go. Prefer an entry a newer event + // of the same record supersedes: the current state still reaches the client // even though the transition history does not. Either way it is a real, // counted loss. auto victim = find_superseded_locked(/*updates_only=*/false); @@ -695,10 +699,9 @@ SSEFaultHandler::resolve_entity_context(const ros2_medkit_msgs::msg::Fault & fau if (fault.reporting_sources.empty()) { return std::nullopt; } - // reporting_sources is a set; debounced faults can carry several co-reporters - // (e.g. node_a and node_b raising the same fault_code). .front() picks the - // lexicographically-first FQN, not a defined owner - any co-reporter's - // rosbag is fetchable, so this remains a valid hint, just not authoritative. + // A record is (fault_code, reporting source), and that source is its owner, + // so reporting_sources has one entry and this is it. The hint addresses the + // record the event is about, not one co-reporter of several. const auto & raw_fqn = fault.reporting_sources.front(); if (raw_fqn.empty()) { return std::nullopt; @@ -737,6 +740,21 @@ SSEFaultHandler::resolve_entity_context(const ros2_medkit_msgs::msg::Fault & fau } } + // A protocol bridge raises its link faults under the COMPONENT's own bare id, + // which is not a ROS FQN and matches no app, so those events carried no + // entity hint at all and a stream consumer could not address the record. An + // external app's bare id needs nothing extra: it has no slash, so the + // last-segment fallback above already looks it up as an app id. + // + // Only an external component claims its bare id as a reporting source, the + // same rule the fault scope applies. Naming a runtime host component would + // point the consumer at an entity whose own fault routes drop the record. + if (entity_id.empty()) { + if (auto component = cache.get_component(raw_fqn); component && component->external.value_or(false)) { + return EntityContext{"components", raw_fqn}; + } + } + if (entity_id.empty()) { RCLCPP_DEBUG(HandlerContext::logger(), "SSE fault event: no entity match for reporting source '%s' (fault_code='%s'); " @@ -745,12 +763,12 @@ SSEFaultHandler::resolve_entity_context(const ros2_medkit_msgs::msg::Fault & fau return std::nullopt; } - // entity_type is hardcoded "apps" because apps are the leaf reporters in - // SOVD - reporting_sources always carries ROS node FQNs which map to apps. - // Components own faults transitively via their hosted apps; consumers can - // walk up the hierarchy via /apps/ -> belongs_to if they need the - // owning component. Manifest-only components without a bound node have no - // FQN match here and fall back to plain discovery - by design. + // entity_type is "apps" on this path because the source resolved to a ROS + // node FQN, and apps are the leaf reporters in SOVD. Components own records + // transitively via their hosted apps, and consumers can walk up the hierarchy via + // /apps/ -> belongs_to if they need the owning component. Manifest-only + // components without a bound node have no FQN match here and fall back to + // plain discovery - by design. return EntityContext{"apps", std::move(entity_id)}; } diff --git a/src/ros2_medkit_gateway/src/http/handlers/trigger_handlers.cpp b/src/ros2_medkit_gateway/src/http/handlers/trigger_handlers.cpp index 59ab2a3db..d6fca9286 100644 --- a/src/ros2_medkit_gateway/src/http/handlers/trigger_handlers.cpp +++ b/src/ros2_medkit_gateway/src/http/handlers/trigger_handlers.cpp @@ -242,21 +242,12 @@ TriggerHandlers::post_trigger(const http::TypedRequest & req, dto::TriggerCreate // (the retry budget is sized against it in gateway_node), and a // genuinely topic-less point surfaces through the retry-expiry warning. auto * pmgr = ctx_.node()->get_plugin_manager(); + // Provider first, vendor route as the fallback - the shared resolution + // the fault-trigger engine's own enumeration uses, so "exists" here and + // "the fetcher could read it" there cannot drift apart. std::optional content; if (pmgr != nullptr) { - if (auto * data_prov = pmgr->get_data_provider_for_entity(entity_id)) { - try { - if (auto result = data_prov->list_data(entity_id)) { - content = result->content; - } - } catch (...) { - // Enumeration is advisory here - a throwing provider must not - // break trigger creation, it just skips the strict check. - } - } - if (!content) { - content = pmgr->fetch_entity_data_via_route(entity_id); - } + content = pmgr->fetch_entity_data_content(entity_id); } // Same liveness guard as the fault-trigger enumeration in gateway_node: // a bridge with a dead PLC link answers 200 with connected=false and an diff --git a/src/ros2_medkit_gateway/src/http/rest_server.cpp b/src/ros2_medkit_gateway/src/http/rest_server.cpp index 5b070aece..41c7d0b10 100644 --- a/src/ros2_medkit_gateway/src/http/rest_server.cpp +++ b/src/ros2_medkit_gateway/src/http/rest_server.cpp @@ -1172,9 +1172,12 @@ void RESTServer::setup_routes() { .requires_role(UserRole::VIEWER) .summary(std::string("Get specific fault for ") + et.singular) .description("Returns fault details including SOVD status, environment data, and rosbag snapshots.") - // 503 when the fault store cannot be read - same branch as the list - // routes, and equally out of the recorder's reach. - .errors({503}) + // 409 (x-medkit-ambiguous-fault) when several sources this entity owns + // report the code and none of their records ranks above the others: + // each is its own record, so the code does not name one. 503 when the + // fault store cannot be read - same branch as the + // list routes, and equally out of the recorder's reach. + .errors({409, 503}) .operation_id(std::string("get") + capitalize(et.singular) + "Fault"); reg.del_alternates( @@ -1189,9 +1192,13 @@ void RESTServer::setup_routes() { .description(std::string("Clears a specific fault for this ") + et.singular + ".") // FaultHandlers::clear_fault -> validate_lock_access("faults"). .lock_guarded() - // 503 when the fault store cannot be read - clear_fault reads the fault - // before clearing it, so it answers the same status the read does. - .errors({503}) + // 409 twice over, from two unrelated causes: lock_guarded() declares the + // locked-entity refusal, and x-medkit-ambiguous-fault answers a code + // that addresses several of this entity's records. Read `vendor_code` + // to tell them apart. 503 when the fault store cannot be read - + // clear_fault reads the record before clearing it, so it answers the + // same status the read does. + .errors({409, 503}) .operation_id(std::string("clear") + capitalize(et.singular) + "Fault"); reg.del(entity_path + "/faults", @@ -1310,6 +1317,11 @@ void RESTServer::setup_routes() { .requires_role(UserRole::VIEWER) .summary(std::string("Download bulk-data file for ") + et.singular) .description("Downloads a bulk-data file (binary content).") + // 409 x-medkit-ambiguous-fault: a rosbag URL carrying a fault code + // (the form that predates recording ids) resolves that code to one + // record the way the fault routes do, and several candidate records + // in this entity's scope name none of them. + .errors({409}) .operation_id(std::string("download") + capitalize(et.singular) + "BulkData"); // Upload: only for apps and components (405 for areas and functions) @@ -2096,6 +2108,8 @@ void RESTServer::setup_routes() { .requires_role(UserRole::VIEWER) .summary("Download bulk-data file for subarea") .description("Downloads a bulk-data file for a subarea.") + // 409 x-medkit-ambiguous-fault, as on the four top-level download routes. + .errors({409}) .operation_id("downloadSubareaBulkData"); // === Nested entities - subcomponents bulk-data === @@ -2130,6 +2144,8 @@ void RESTServer::setup_routes() { .requires_role(UserRole::VIEWER) .summary("Download bulk-data file for subcomponent") .description("Downloads a bulk-data file for a subcomponent.") + // 409 x-medkit-ambiguous-fault, as on the four top-level download routes. + .errors({409}) .operation_id("downloadSubcomponentBulkData"); // === Global faults === diff --git a/src/ros2_medkit_gateway/src/plugins/plugin_manager.cpp b/src/ros2_medkit_gateway/src/plugins/plugin_manager.cpp index 47b8e7036..7fa4d00bb 100644 --- a/src/ros2_medkit_gateway/src/plugins/plugin_manager.cpp +++ b/src/ros2_medkit_gateway/src/plugins/plugin_manager.cpp @@ -431,6 +431,25 @@ bool PluginManager::has_entity_data_route(const std::string & entity_id) { return false; } +std::optional PluginManager::fetch_entity_data_content(const std::string & entity_id) { + // Rationale is on the declaration: provider first, vendor route only when the + // owner exposes no provider for this entity. + if (auto * data_prov = get_data_provider_for_entity(entity_id)) { + try { + if (auto result = data_prov->list_data(entity_id)) { + return result->content; + } + } catch (const std::exception & e) { + RCLCPP_WARN(rclcpp::get_logger("plugin_manager"), "DataProvider threw for entity '%s': %s. Trying its route.", + entity_id.c_str(), e.what()); + } catch (...) { + RCLCPP_WARN(rclcpp::get_logger("plugin_manager"), "DataProvider threw for entity '%s'. Trying its route.", + entity_id.c_str()); + } + } + return fetch_entity_data_via_route(entity_id); +} + std::optional PluginManager::fetch_entity_data_via_route(const std::string & entity_id, const std::string & item) { const std::string full_path = diff --git a/src/ros2_medkit_gateway/src/ros2/conversions/fault_msg_conversions.cpp b/src/ros2_medkit_gateway/src/ros2/conversions/fault_msg_conversions.cpp index 9d6833187..8c546fd98 100644 --- a/src/ros2_medkit_gateway/src/ros2/conversions/fault_msg_conversions.cpp +++ b/src/ros2_medkit_gateway/src/ros2/conversions/fault_msg_conversions.cpp @@ -42,6 +42,12 @@ nlohmann::json fault_to_json(const ros2_medkit_msgs::msg::Fault & fault) { j["occurrence_count"] = fault.occurrence_count; j["status"] = fault.status; j["reporting_sources"] = fault.reporting_sources; + // The owner of the record, next to the list it is the single entry of. It is + // what the per-record routes send back as source_id, so a client never has to + // reach into an array to address the record it is looking at. + if (!fault.reporting_sources.empty()) { + j["source_id"] = fault.reporting_sources.front(); + } switch (fault.severity) { case ros2_medkit_msgs::msg::Fault::SEVERITY_INFO: diff --git a/src/ros2_medkit_gateway/src/ros2/transports/ros2_fault_service_transport.cpp b/src/ros2_medkit_gateway/src/ros2/transports/ros2_fault_service_transport.cpp index 935365a19..eddd863ac 100644 --- a/src/ros2_medkit_gateway/src/ros2/transports/ros2_fault_service_transport.cpp +++ b/src/ros2_medkit_gateway/src/ros2/transports/ros2_fault_service_transport.cpp @@ -256,10 +256,14 @@ FaultResult Ros2FaultServiceTransport::list_faults(const std::string & source_id if (include_muted && !response->muted_faults.empty()) { json muted_array = json::array(); for (const auto & muted : response->muted_faults) { + // source_id is what makes the entry addressable: a muted record is one + // record, and its code alone names as many as there are sources + // reporting it. muted_array.push_back({{"fault_code", muted.fault_code}, {"root_cause_code", muted.root_cause_code}, {"rule_id", muted.rule_id}, - {"delay_ms", muted.delay_ms}}); + {"delay_ms", muted.delay_ms}, + {"source_id", muted.source_id}}); } result.data["muted_faults"] = muted_array; } @@ -294,6 +298,11 @@ FaultWithEnvJsonResult Ros2FaultServiceTransport::get_fault_with_env(const std:: auto request = std::make_shared(); request->fault_code = fault_code; + // The record is (fault_code, source_id), so the owner goes in the request and + // the fault manager resolves the record. Filtering the answer here instead + // would mean asking for an arbitrary record and then hoping it is the one the + // caller meant. + request->source_id = source_id; auto response = invoke_fault_service( get_fault_client_, request, executor_, executor_mutex_, std::chrono::duration(service_timeout_sec_), @@ -310,22 +319,6 @@ FaultWithEnvJsonResult Ros2FaultServiceTransport::get_fault_with_env(const std:: return result; } - // Verify source_id if provided (prefix match against any reporting source). - if (!source_id.empty()) { - bool matches = false; - for (const auto & src : response->fault.reporting_sources) { - if (src.rfind(source_id, 0) == 0) { - matches = true; - break; - } - } - if (!matches) { - result.success = false; - result.error_message = "Fault not found for source: " + source_id; - return result; - } - } - result.success = true; result.data = {{"fault", conversions::fault_to_json(response->fault)}, {"environment_data", conversions::environment_data_to_json(response->environment_data)}}; @@ -348,11 +341,13 @@ FaultResult Ros2FaultServiceTransport::get_fault(const std::string & fault_code, return result; } -FaultResult Ros2FaultServiceTransport::clear_fault(const std::string & fault_code, bool skip_correlation_auto_clear) { +FaultResult Ros2FaultServiceTransport::clear_fault(const std::string & fault_code, const std::string & source_id, + bool skip_correlation_auto_clear) { FaultResult result; auto request = std::make_shared(); request->fault_code = fault_code; + request->source_id = source_id; request->skip_correlation_auto_clear = skip_correlation_auto_clear; auto response = invoke_fault_service( @@ -377,11 +372,13 @@ FaultResult Ros2FaultServiceTransport::clear_fault(const std::string & fault_cod return result; } -FaultResult Ros2FaultServiceTransport::get_snapshots(const std::string & fault_code, const std::string & topic) { +FaultResult Ros2FaultServiceTransport::get_snapshots(const std::string & fault_code, const std::string & source_id, + const std::string & topic) { FaultResult result; auto request = std::make_shared(); request->fault_code = fault_code; + request->source_id = source_id; request->topic = topic; auto response = invoke_fault_service( @@ -412,16 +409,18 @@ FaultResult Ros2FaultServiceTransport::get_snapshots(const std::string & fault_c return result; } -FaultResult Ros2FaultServiceTransport::get_rosbag(const std::string & fault_code) { +FaultResult Ros2FaultServiceTransport::get_rosbag(const std::string & id, const std::string & source_id) { FaultResult result; // The parameter is the bulk-data id, which is a recording id. It is sent in BOTH // fields: the fault manager prefers recording_id and falls back to fault_code, // which is what keeps a pre-#620 URL (and every existing .test.py that calls this - // service with a fault code) working unchanged. + // service with a fault code) working unchanged. source_id scopes that fallback + // to one record. The recording_id path ignores it. auto request = std::make_shared(); - request->recording_id = fault_code; - request->fault_code = fault_code; + request->recording_id = id; + request->fault_code = id; + request->source_id = source_id; auto response = invoke_fault_service( get_rosbag_client_, request, executor_, executor_mutex_, std::chrono::duration(service_timeout_sec_), diff --git a/src/ros2_medkit_gateway/test/test_bulkdata_handlers.cpp b/src/ros2_medkit_gateway/test/test_bulkdata_handlers.cpp index c98d65ea1..c990322e2 100644 --- a/src/ros2_medkit_gateway/test/test_bulkdata_handlers.cpp +++ b/src/ros2_medkit_gateway/test/test_bulkdata_handlers.cpp @@ -111,15 +111,28 @@ TEST_F(BulkDataHandlersTest, RecordingIdToleratesTrailingSlashAndEmptyPath) { namespace { -json rosbag_row(const std::string & fault_code, const std::string & recording_id, uint64_t size_bytes = 1024) { - return json{{"fault_code", fault_code}, {"recording_id", recording_id}, {"file_path", "/var/bags/" + recording_id}, - {"format", "mcap"}, {"duration_sec", 5.0}, {"size_bytes", size_bytes}}; +/// The listing stamps each row with the reporting source it was listed under, +/// which is the owner of the record the recording belongs to. +json rosbag_row(const std::string & fault_code, const std::string & recording_id, uint64_t size_bytes = 1024, + const std::string & owner = "app_a") { + return json{{"fault_code", fault_code}, + {"recording_id", recording_id}, + {"file_path", "/var/bags/" + recording_id}, + {"format", "mcap"}, + {"duration_sec", 5.0}, + {"size_bytes", size_bytes}, + {"owner", owner}}; } json fault_at(double first_occurred) { return json{{"first_occurred", first_occurred}}; } +/// Key one record for the date lookup the same way the handler does. +std::string record_key(const std::string & owner, const std::string & fault_code) { + return handlers::detail::record_map_key(owner, fault_code); +} + /// A row carrying the recording's own timestamp, which is what the fault manager /// sends now. json rosbag_row_made_at(const std::string & fault_code, const std::string & recording_id, int64_t created_at_ns) { @@ -179,8 +192,8 @@ TEST_F(BulkDataHandlersTest, ARecordingIsDatedByTheEarliestFaultOfItsBurst) { // Downstream faults confirm after the root cause, and the recording covers // the whole burst, so the earliest is the honest creation date. const std::vector rows{rosbag_row("DOWNSTREAM", "fault_ROOT_9"), rosbag_row("ROOT", "fault_ROOT_9")}; - const std::unordered_map faults{{"DOWNSTREAM", fault_at(1700000900.0)}, - {"ROOT", fault_at(1700000000.0)}}; + const std::unordered_map faults{{record_key("app_a", "DOWNSTREAM"), fault_at(1700000900.0)}, + {record_key("app_a", "ROOT"), fault_at(1700000000.0)}}; const auto descriptors = handlers::detail::fold_rosbag_rows_into_descriptors(rows, faults); ASSERT_EQ(descriptors.size(), 1u); @@ -193,7 +206,7 @@ TEST_F(BulkDataHandlersTest, EachRecordingOfOneFaultIsDatedByItsOwnCapture) { // is exactly what tells the occurrences apart. const std::vector rows{rosbag_row_made_at("FLAP", "fault_FLAP_2", int64_t{1700000900} * 1'000'000'000), rosbag_row_made_at("FLAP", "fault_FLAP_1", int64_t{1700000000} * 1'000'000'000)}; - const std::unordered_map faults{{"FLAP", fault_at(1700000000.0)}}; + const std::unordered_map faults{{record_key("app_a", "FLAP"), fault_at(1700000000.0)}}; const auto descriptors = handlers::detail::fold_rosbag_rows_into_descriptors(rows, faults); ASSERT_EQ(descriptors.size(), 2u); @@ -202,6 +215,23 @@ TEST_F(BulkDataHandlersTest, EachRecordingOfOneFaultIsDatedByItsOwnCapture) { EXPECT_NE(descriptors[0].creation_date, descriptors[1].creation_date); } +// Two sources reporting one code are two records with their own first_occurred. +// Keying the date lookup on the code alone let whichever record was listed last +// date the other owner's recordings. +TEST_F(BulkDataHandlersTest, ARecordingIsDatedByItsOwnOwnersRecord) { + const std::vector rows{rosbag_row("SHARED_CODE", "rec_a", 1024, "app_a"), + rosbag_row("SHARED_CODE", "rec_b", 1024, "app_b")}; + const std::unordered_map faults{{record_key("app_a", "SHARED_CODE"), fault_at(1700000000.0)}, + {record_key("app_b", "SHARED_CODE"), fault_at(1700009999.0)}}; + + const auto descriptors = handlers::detail::fold_rosbag_rows_into_descriptors(rows, faults); + + ASSERT_EQ(descriptors.size(), 2u); + EXPECT_EQ(descriptors[0].creation_date, format_timestamp_ns(int64_t{1700000000} * 1'000'000'000)); + EXPECT_EQ(descriptors[1].creation_date, format_timestamp_ns(int64_t{1700009999} * 1'000'000'000)) + << "app_b's recording was dated from app_a's record of the same code"; +} + TEST_F(BulkDataHandlersTest, AnAcknowledgedFaultsRecordingsKeepTheirRealDate) { // list_faults excludes cleared faults by default, so an acknowledged fault is // absent from the map while its rows are still listed. Reading the date off the @@ -651,9 +681,11 @@ TEST_F(BulkDataSourceFiltersTest, FunctionWithComponentHostResolvesComponentApps // get_fault(code, source) semantics) would let app id "plc" claim the bag of // "plc_line1". // === Download authorization tests === -// A recording is shared by a whole burst, so ownership is the union over its -// attached faults. The scope matcher itself is unchanged and pinned below; what -// is new is which codes get fed to it. +// A recording is shared by a whole burst, so ownership is the union over the +// records it is attached to, each a (fault code, owner) pair. The scope matcher +// itself is unchanged and pinned below. The attached codes say which of the +// entity's records to ask about, and the owner's own rosbag rows say whether it +// holds the recording. TEST_F(BulkDataSourceFiltersTest, AttachedFaultCodesComeFromTheRecordingNotTheUrl) { const nlohmann::json rosbag = {{"file_path", "/var/bags/fault_ROOT_1"}, @@ -683,6 +715,44 @@ TEST_F(BulkDataSourceFiltersTest, AttachedFaultCodesFallBackOnAnEmptyOrMalformed EXPECT_EQ(handlers::detail::rosbag_attached_fault_codes(not_an_array, "X"), (std::vector{"X"})); } +// One owner's ListRosbags rows prove its records are attached to the recording +// only when a row names that recording. Another recording of the same code is +// another owner's, or another occurrence, and proves nothing. +TEST_F(BulkDataSourceFiltersTest, AnOwnersRowsHoldARecordingOnlyWhenARowNamesIt) { + const nlohmann::json served = {{"file_path", "/var/bags/rec_pump"}, {"recording_id", "rec_pump"}}; + const nlohmann::json tank_rows = nlohmann::json::array( + {nlohmann::json{{"fault_code", "SHARED"}, {"recording_id", "rec_tank"}, {"file_path", "/var/bags/rec_tank"}}}); + const nlohmann::json pump_rows = nlohmann::json::array( + {nlohmann::json{{"fault_code", "SHARED"}, {"recording_id", "rec_pump"}, {"file_path", "/var/bags/rec_pump"}}}); + + EXPECT_FALSE(handlers::detail::rosbag_rows_hold_recording(tank_rows, served)) + << "tank's recording of the same code is not pump's recording"; + EXPECT_TRUE(handlers::detail::rosbag_rows_hold_recording(pump_rows, served)); + EXPECT_FALSE(handlers::detail::rosbag_rows_hold_recording(nlohmann::json::array(), served)); + EXPECT_FALSE(handlers::detail::rosbag_rows_hold_recording(nlohmann::json::object(), served)); +} + +// A peer that predates recording ids names a bag only by its path, on either +// side. The path then identifies the recording. Two ids that differ never match +// through their paths. +TEST_F(BulkDataSourceFiltersTest, ARecordingIsMatchedByItsPathWhenEitherSideHasNoId) { + const nlohmann::json row_without_id = + nlohmann::json::array({nlohmann::json{{"fault_code", "SHARED"}, {"file_path", "/var/bags/rec_pump"}}}); + const nlohmann::json served_with_id = {{"file_path", "/var/bags/rec_pump"}, {"recording_id", "rec_pump"}}; + const nlohmann::json served_without_id = {{"file_path", "/var/bags/rec_pump"}}; + const nlohmann::json row_with_id = nlohmann::json::array( + {nlohmann::json{{"fault_code", "SHARED"}, {"recording_id", "rec_pump"}, {"file_path", "/var/bags/rec_pump"}}}); + + EXPECT_TRUE(handlers::detail::rosbag_rows_hold_recording(row_without_id, served_with_id)); + EXPECT_TRUE(handlers::detail::rosbag_rows_hold_recording(row_with_id, served_without_id)); + EXPECT_FALSE(handlers::detail::rosbag_rows_hold_recording(row_without_id, {{"file_path", "/var/bags/rec_tank"}})); + + const nlohmann::json same_path_other_id = nlohmann::json::array( + {nlohmann::json{{"fault_code", "SHARED"}, {"recording_id", "rec_other"}, {"file_path", "/var/bags/rec_pump"}}}); + EXPECT_FALSE(handlers::detail::rosbag_rows_hold_recording(same_path_other_id, served_with_id)) + << "two recording ids are two recordings"; +} + TEST_F(BulkDataSourceFiltersTest, AFaultCodeUrlIsRecognisedAsTheCompatibilityPath) { // The segment named a fault; the answer is that fault's newest recording, whose // id is something else. Authorizing on the union alone would 200 a code the diff --git a/src/ros2_medkit_gateway/test/test_entity_freeze_frame_capture.cpp b/src/ros2_medkit_gateway/test/test_entity_freeze_frame_capture.cpp index b9a44b2e0..c7d6fab45 100644 --- a/src/ros2_medkit_gateway/test/test_entity_freeze_frame_capture.cpp +++ b/src/ros2_medkit_gateway/test/test_entity_freeze_frame_capture.cpp @@ -181,7 +181,7 @@ class EntityFreezeFrameCaptureTest : public ::testing::Test { const auto deadline = std::chrono::steady_clock::now() + 5s; while (std::chrono::steady_clock::now() < deadline) { publisher_->publish(event); - if (!capture.frames_for(event.fault.fault_code).empty()) { + if (!capture.frames_for(event.fault.fault_code, event.fault.reporting_sources.front()).empty()) { return true; } std::this_thread::sleep_for(20ms); @@ -208,7 +208,7 @@ TEST_F(EntityFreezeFrameCaptureTest, ConfirmedPluginFaultCapturesEntityValues) { ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_OVERPRESSURE", {"plc_app"}))); - auto frames = capture.frames_for("PLC_OVERPRESSURE"); + auto frames = capture.frames_for("PLC_OVERPRESSURE", "plc_app"); ASSERT_EQ(frames.size(), 1u); EXPECT_EQ(frames[0].entity_id, "plc_app"); EXPECT_DOUBLE_EQ(frames[0].values.value("temperature", 0.0), 42.5); @@ -230,7 +230,7 @@ TEST_F(EntityFreezeFrameCaptureTest, NonPluginSourceCapturesNothing) { publisher_->publish(event); std::this_thread::sleep_for(10ms); } - EXPECT_TRUE(capture.frames_for("ROS_FAULT").empty()); + EXPECT_TRUE(capture.frames_for("ROS_FAULT", "/powertrain/engine/temp_sensor").empty()); } /// @verifies REQ_INTEROP_088 @@ -246,7 +246,7 @@ TEST_F(EntityFreezeFrameCaptureTest, NonConfirmedEventsAreIgnored) { publisher_->publish(event); std::this_thread::sleep_for(10ms); } - EXPECT_TRUE(capture.frames_for("PLC_UPDATED_ONLY").empty()); + EXPECT_TRUE(capture.frames_for("PLC_UPDATED_ONLY", "plc_app").empty()); } /// @verifies REQ_INTEROP_088 @@ -268,10 +268,10 @@ TEST_F(EntityFreezeFrameCaptureTest, StandingFaultsAreFramedWithoutAnyEvent) { // No publish at all - the frame must appear from the catch-up alone. const auto deadline = std::chrono::steady_clock::now() + 15s; - while (capture.frames_for("PLC_STANDING").empty() && std::chrono::steady_clock::now() < deadline) { + while (capture.frames_for("PLC_STANDING", "route_plc_app").empty() && std::chrono::steady_clock::now() < deadline) { std::this_thread::sleep_for(20ms); } - const auto frames = capture.frames_for("PLC_STANDING"); + const auto frames = capture.frames_for("PLC_STANDING", "route_plc_app"); ASSERT_EQ(frames.size(), 1u); EXPECT_EQ(frames[0].entity_id, "route_plc_app"); EXPECT_DOUBLE_EQ(frames[0].values.value("level", 0.0), 7.0); @@ -279,29 +279,24 @@ TEST_F(EntityFreezeFrameCaptureTest, StandingFaultsAreFramedWithoutAnyEvent) { } /// @verifies REQ_INTEROP_088 -TEST_F(EntityFreezeFrameCaptureTest, QueuedLiveConfirmDedupesStandingCatchUp) { - // Both paths carry the same fault code and the confirm is already queued - // when the catch-up snapshots the queue: the catch-up must skip that code - // entirely (one plugin read per confirm, no double capture) and the drain - // loop's live frame is the one that lands. +TEST_F(EntityFreezeFrameCaptureTest, QueuedLiveConfirmDedupesTheSameStandingRecord) { + // The same record - one code, one source - arrives on both paths, and the + // confirm is already queued when the catch-up snapshots the queue: the + // catch-up must skip that record (one plugin read per confirm, no double + // capture) and the drain loop's live frame is the one that lands. std::atomic published{false}; - std::atomic standing_sampled{false}; + std::atomic route_reads{0}; EntityFreezeFrameCapture capture( node_.get(), *sub_exec_, [](const std::string &) -> DataProvider * { return nullptr; }, - [&standing_sampled](const std::string & entity_id) -> std::optional { - double level = 0.0; - if (entity_id == "route_live_app") { - level = 99.0; - } else if (entity_id == "route_standing_app") { - level = 1.0; - standing_sampled.store(true); - } else { + [&route_reads](const std::string & entity_id) -> std::optional { + if (entity_id != "route_live_app") { return std::nullopt; } - return json{{"connected", true}, {"items", json::array({{{"name", "level"}, {"value", level}}})}}; + route_reads.fetch_add(1); + return json{{"connected", true}, {"items", json::array({{{"name", "level"}, {"value", 99.0}}})}}; }, 256, [&published](const std::function &) -> std::vector { @@ -312,7 +307,7 @@ TEST_F(EntityFreezeFrameCaptureTest, QueuedLiveConfirmDedupesStandingCatchUp) { std::this_thread::sleep_for(10ms); } std::this_thread::sleep_for(500ms); // let the event reach the queue - return {{"PLC_BOTH_PATHS", {"route_standing_app"}}}; + return {{"PLC_BOTH_PATHS", {"route_live_app"}}}; }); ASSERT_TRUE(wait_for_match()); @@ -323,22 +318,114 @@ TEST_F(EntityFreezeFrameCaptureTest, QueuedLiveConfirmDedupesStandingCatchUp) { const auto deadline = std::chrono::steady_clock::now() + 10s; std::vector frames; while (std::chrono::steady_clock::now() < deadline) { - frames = capture.frames_for("PLC_BOTH_PATHS"); + frames = capture.frames_for("PLC_BOTH_PATHS", "route_live_app"); if (!frames.empty()) { break; } std::this_thread::sleep_for(20ms); } - // The queued confirm dedupes the catch-up: the standing entity was never - // sampled and the only frame is the live one. - EXPECT_FALSE(standing_sampled.load()); + // The queued confirm dedupes the catch-up: the record was read once and + // holds the live frame. + std::this_thread::sleep_for(300ms); // would be enough for a second read + EXPECT_EQ(route_reads.load(), 1); ASSERT_EQ(frames.size(), 1u); EXPECT_EQ(frames[0].entity_id, "route_live_app"); EXPECT_DOUBLE_EQ(frames[0].values.value("level", 0.0), 99.0); EXPECT_FALSE(frames[0].startup_catchup); } +/// A queued confirm dedupes its own record, not the code. Another source +/// standing on the same code is a different record with its own frame, and +/// dropping it on the code alone left that record's detail page blank. +/// @verifies REQ_INTEROP_088 +TEST_F(EntityFreezeFrameCaptureTest, QueuedConfirmDoesNotDedupeAnotherOwnerOfTheSameCode) { + std::atomic published{false}; + EntityFreezeFrameCapture capture( + node_.get(), *sub_exec_, + [](const std::string &) -> DataProvider * { + return nullptr; + }, + [](const std::string & entity_id) -> std::optional { + double level = 0.0; + if (entity_id == "route_live_app") { + level = 99.0; + } else if (entity_id == "route_standing_app") { + level = 1.0; + } else { + return std::nullopt; + } + return json{{"connected", true}, {"items", json::array({{{"name", "level"}, {"value", level}}})}}; + }, + 256, + [&published](const std::function &) -> std::vector { + const auto deadline = std::chrono::steady_clock::now() + 5s; + while (!published.load() && std::chrono::steady_clock::now() < deadline) { + std::this_thread::sleep_for(10ms); + } + std::this_thread::sleep_for(500ms); // let the event reach the queue + return {{"PLC_TWO_OWNERS", {"route_standing_app"}}}; + }); + + ASSERT_TRUE(wait_for_match()); + publisher_->publish(make_confirmed_event("PLC_TWO_OWNERS", {"route_live_app"})); + published.store(true); + + const auto deadline = std::chrono::steady_clock::now() + 10s; + std::vector live_frames; + std::vector standing_frames; + while (std::chrono::steady_clock::now() < deadline) { + live_frames = capture.frames_for("PLC_TWO_OWNERS", "route_live_app"); + standing_frames = capture.frames_for("PLC_TWO_OWNERS", "route_standing_app"); + if (!live_frames.empty() && !standing_frames.empty()) { + break; + } + std::this_thread::sleep_for(20ms); + } + + ASSERT_EQ(live_frames.size(), 1u) << "the live record lost its frame"; + EXPECT_DOUBLE_EQ(live_frames[0].values.value("level", 0.0), 99.0); + ASSERT_EQ(standing_frames.size(), 1u) << "the standing record of the same code was dropped as a duplicate"; + EXPECT_DOUBLE_EQ(standing_frames[0].values.value("level", 0.0), 1.0); + EXPECT_TRUE(standing_frames[0].startup_catchup); +} + +/// Two sources confirming one code are two records, each keeping its own +/// values. Under a code-only key the second confirmation overwrote the first +/// source's frames and served them on the first source's detail page. +/// @verifies REQ_INTEROP_088 +TEST_F(EntityFreezeFrameCaptureTest, TwoOwnersOfOneCodeKeepSeparateFrames) { + EntityFreezeFrameCapture capture( + node_.get(), *sub_exec_, + [](const std::string &) -> DataProvider * { + return nullptr; + }, + [](const std::string & entity_id) -> std::optional { + double level = 0.0; + if (entity_id == "line_a") { + level = 11.0; + } else if (entity_id == "line_b") { + level = 22.0; + } else { + return std::nullopt; + } + return json{{"connected", true}, {"items", json::array({{{"name", "level"}, {"value", level}}})}}; + }); + + ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("SHARED_CODE", {"line_a"}))); + ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("SHARED_CODE", {"line_b"}))); + + const auto a = capture.frames_for("SHARED_CODE", "line_a"); + const auto b = capture.frames_for("SHARED_CODE", "line_b"); + + ASSERT_EQ(a.size(), 1u) << "line_a's record lost its frame to line_b's confirmation"; + EXPECT_EQ(a[0].entity_id, "line_a"); + EXPECT_DOUBLE_EQ(a[0].values.value("level", 0.0), 11.0); + ASSERT_EQ(b.size(), 1u); + EXPECT_EQ(b[0].entity_id, "line_b"); + EXPECT_DOUBLE_EQ(b[0].values.value("level", 0.0), 22.0); +} + /// @verifies REQ_INTEROP_088 TEST_F(EntityFreezeFrameCaptureTest, ThrowingListerDoesNotKillCaptureThread) { // The lister runs on the capture thread's entry path: an escaping exception @@ -355,7 +442,7 @@ TEST_F(EntityFreezeFrameCaptureTest, ThrowingListerDoesNotKillCaptureThread) { }); ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_AFTER_THROW", {"plc_app"}))); - EXPECT_FALSE(capture.frames_for("PLC_AFTER_THROW").empty()); + EXPECT_FALSE(capture.frames_for("PLC_AFTER_THROW", "plc_app").empty()); } /// @verifies REQ_INTEROP_088 @@ -452,14 +539,14 @@ TEST_F(EntityFreezeFrameCaptureTest, CatchUpCapsStoredFramesAtMaxFaults) { }); const auto deadline = std::chrono::steady_clock::now() + 15s; - while (capture.frames_for("PLC_CAP_B").empty() && std::chrono::steady_clock::now() < deadline) { + while (capture.frames_for("PLC_CAP_B", "route_b").empty() && std::chrono::steady_clock::now() < deadline) { std::this_thread::sleep_for(20ms); } std::this_thread::sleep_for(300ms); // would be enough for an uncapped C capture - EXPECT_TRUE(capture.frames_for("PLC_CAP_NO_DATA").empty()); - EXPECT_FALSE(capture.frames_for("PLC_CAP_A").empty()); // not evicted by C - EXPECT_FALSE(capture.frames_for("PLC_CAP_B").empty()); - EXPECT_TRUE(capture.frames_for("PLC_CAP_C").empty()); // past the cap + EXPECT_TRUE(capture.frames_for("PLC_CAP_NO_DATA", "route_no_data_app").empty()); + EXPECT_FALSE(capture.frames_for("PLC_CAP_A", "route_a").empty()); // not evicted by C + EXPECT_FALSE(capture.frames_for("PLC_CAP_B", "route_b").empty()); + EXPECT_TRUE(capture.frames_for("PLC_CAP_C", "route_c").empty()); // past the cap } /// @verifies REQ_INTEROP_088 @@ -469,7 +556,7 @@ TEST_F(EntityFreezeFrameCaptureTest, FramesRetainedAcrossClearAndOverwrittenOnRe }); ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_CYCLING", {"plc_app"}))); - ASSERT_DOUBLE_EQ(capture.frames_for("PLC_CYCLING")[0].values.value("temperature", 0.0), 42.5); + ASSERT_DOUBLE_EQ(capture.frames_for("PLC_CYCLING", "plc_app")[0].values.value("temperature", 0.0), 42.5); // Clearing retains the confirmed-state record, mirroring the fault_manager's // freeze-frame retention across clear_fault. @@ -479,7 +566,7 @@ TEST_F(EntityFreezeFrameCaptureTest, FramesRetainedAcrossClearAndOverwrittenOnRe publisher_->publish(cleared); std::this_thread::sleep_for(10ms); } - auto retained = capture.frames_for("PLC_CYCLING"); + auto retained = capture.frames_for("PLC_CYCLING", "plc_app"); ASSERT_EQ(retained.size(), 1u); EXPECT_DOUBLE_EQ(retained[0].values.value("temperature", 0.0), 42.5); @@ -492,7 +579,7 @@ TEST_F(EntityFreezeFrameCaptureTest, FramesRetainedAcrossClearAndOverwrittenOnRe while (std::chrono::steady_clock::now() < deadline && !overwritten) { publisher_->publish(reconfirm); std::this_thread::sleep_for(20ms); - auto latest = capture.frames_for("PLC_CYCLING"); + auto latest = capture.frames_for("PLC_CYCLING", "plc_app"); overwritten = !latest.empty() && std::abs(latest[0].values.value("temperature", 0.0) - 99.0) < 1e-9; } EXPECT_TRUE(overwritten); @@ -518,7 +605,7 @@ TEST_F(EntityFreezeFrameCaptureTest, RouteFallbackCapturesWhenPluginHasNoDataPro ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_ROUTE_LEVEL_HIGH", {"route_plc_app"}))); - auto frames = capture.frames_for("PLC_ROUTE_LEVEL_HIGH"); + auto frames = capture.frames_for("PLC_ROUTE_LEVEL_HIGH", "route_plc_app"); ASSERT_EQ(frames.size(), 1u); EXPECT_EQ(frames[0].entity_id, "route_plc_app"); EXPECT_DOUBLE_EQ(frames[0].values.value("level", 0.0), 87.5); @@ -546,7 +633,7 @@ TEST_F(EntityFreezeFrameCaptureTest, RouteFallbackSkipsDisconnectedPlcWithNoValu publisher_->publish(event); std::this_thread::sleep_for(10ms); } - EXPECT_TRUE(capture.frames_for("PLC_DISCONNECTED").empty()); + EXPECT_TRUE(capture.frames_for("PLC_DISCONNECTED", "route_plc_app").empty()); } /// @verifies REQ_INTEROP_088 @@ -567,7 +654,7 @@ TEST_F(EntityFreezeFrameCaptureTest, DisconnectedEntityWithLastKnownValuesIsCapt ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_COMMS_LOST", {"route_plc_app"}))); - const auto frames = capture.frames_for("PLC_COMMS_LOST"); + const auto frames = capture.frames_for("PLC_COMMS_LOST", "route_plc_app"); ASSERT_EQ(frames.size(), 1u); EXPECT_EQ(frames[0].values["level"], 42.0); ASSERT_TRUE(frames[0].connected.has_value()); @@ -590,7 +677,7 @@ TEST_F(EntityFreezeFrameCaptureTest, DataProviderWinsOverRouteFallback) { ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_PROVIDER_FIRST", {"plc_app"}))); - auto frames = capture.frames_for("PLC_PROVIDER_FIRST"); + auto frames = capture.frames_for("PLC_PROVIDER_FIRST", "plc_app"); ASSERT_EQ(frames.size(), 1u); EXPECT_DOUBLE_EQ(frames[0].values.value("temperature", 0.0), 42.5); // provider values, not route EXPECT_FALSE(fetcher_called.load()); @@ -620,7 +707,7 @@ TEST_F(EntityFreezeFrameCaptureTest, DataProviderWithoutLiveValuesCapturesNothin publisher_->publish(event); std::this_thread::sleep_for(10ms); } - EXPECT_TRUE(capture.frames_for("PLC_NO_LIVE_DATA").empty()); + EXPECT_TRUE(capture.frames_for("PLC_NO_LIVE_DATA", "cold_app").empty()); } /// @verifies REQ_INTEROP_088 @@ -636,7 +723,7 @@ TEST_F(EntityFreezeFrameCaptureTest, DisconnectedDataProviderWithLastKnownValues ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_PROVIDER_COMMS_LOST", {"down_app"}))); - const auto frames = capture.frames_for("PLC_PROVIDER_COMMS_LOST"); + const auto frames = capture.frames_for("PLC_PROVIDER_COMMS_LOST", "down_app"); ASSERT_EQ(frames.size(), 1u); EXPECT_EQ(frames[0].values["level"], 5.5); ASSERT_TRUE(frames[0].connected.has_value()); @@ -657,9 +744,9 @@ TEST_F(EntityFreezeFrameCaptureTest, OldestFaultEvictedPastMaxFaults) { ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_EVICT_B", {"plc_app"}))); ASSERT_TRUE(publish_and_wait(capture, make_confirmed_event("PLC_EVICT_C", {"plc_app"}))); - EXPECT_TRUE(capture.frames_for("PLC_EVICT_A").empty()); // FIFO-evicted - EXPECT_FALSE(capture.frames_for("PLC_EVICT_B").empty()); - EXPECT_FALSE(capture.frames_for("PLC_EVICT_C").empty()); + EXPECT_TRUE(capture.frames_for("PLC_EVICT_A", "plc_app").empty()); // FIFO-evicted + EXPECT_FALSE(capture.frames_for("PLC_EVICT_B", "plc_app").empty()); + EXPECT_FALSE(capture.frames_for("PLC_EVICT_C", "plc_app").empty()); } TEST(ContentHasLiveData, GatesOnItemsNotOnTheLinkFlag) { diff --git a/src/ros2_medkit_gateway/test/test_fault_handlers.cpp b/src/ros2_medkit_gateway/test/test_fault_handlers.cpp index 9c7f08a91..208359b18 100644 --- a/src/ros2_medkit_gateway/test/test_fault_handlers.cpp +++ b/src/ros2_medkit_gateway/test/test_fault_handlers.cpp @@ -16,6 +16,7 @@ #include +#include "ros2_medkit_gateway/core/http/error_codes.hpp" #include "ros2_medkit_gateway/dto/faults.hpp" #include "ros2_medkit_gateway/dto/json_reader.hpp" #include "ros2_medkit_gateway/dto/json_writer.hpp" @@ -381,11 +382,16 @@ TEST_F(FaultHandlersTest, BuildSovdFaultResponsePrimaryValueExtraction) { } } +// A record is (fault_code, owner), so the detail names exactly one reporting +// source and repeats it as x-medkit.owner, which is the value the per-record +// routes address the record by. This test used to pin one record +// carrying three sources. Three sources reporting one code are three records +// now, and each detail response describes one of them. // @verifies REQ_INTEROP_013 -TEST_F(FaultHandlersTest, BuildSovdFaultResponseMultipleSources) { +TEST_F(FaultHandlersTest, BuildSovdFaultResponseNamesTheOwningSource) { ros2_medkit_msgs::msg::Fault fault; - fault.fault_code = "MULTI_SOURCE_FAULT"; - fault.reporting_sources = {"/perception/lidar", "/perception/camera", "/control/motor"}; + fault.fault_code = "SHARED_CODE"; + fault.reporting_sources = {"/perception/lidar"}; ros2_medkit_msgs::msg::EnvironmentData env_data; @@ -393,10 +399,36 @@ TEST_F(FaultHandlersTest, BuildSovdFaultResponseMultipleSources) { to_json(FaultHandlers::build_sovd_fault_response(fault_json(fault), env_json(env_data), "/apps/test")); auto sources = response["x-medkit"]["reporting_sources"]; - ASSERT_EQ(sources.size(), 3); + ASSERT_EQ(sources.size(), 1); EXPECT_EQ(sources[0], "/perception/lidar"); - EXPECT_EQ(sources[1], "/perception/camera"); - EXPECT_EQ(sources[2], "/control/motor"); + ASSERT_TRUE(response["x-medkit"].contains("owner")); + EXPECT_EQ(response["x-medkit"]["owner"], "/perception/lidar"); + // Named `owner`, never `source_id`: a fault LIST's x-medkit.source_id is the + // addressed entity's namespace path, so one key with two meanings inside one + // API is the thing this avoids. + EXPECT_FALSE(response["x-medkit"].contains("source_id")) + << "the detail must not reuse the list's source_id key for the record owner"; +} + +// The other owner of the same code is a different record, and its detail says +// so: same code, different source_id. +// @verifies REQ_INTEROP_013 +TEST_F(FaultHandlersTest, BuildSovdFaultResponseSeparatesOwnersOfOneCode) { + ros2_medkit_msgs::msg::Fault first; + first.fault_code = "SHARED_CODE"; + first.reporting_sources = {"/perception/lidar"}; + ros2_medkit_msgs::msg::Fault second; + second.fault_code = "SHARED_CODE"; + second.reporting_sources = {"/control/motor"}; + + ros2_medkit_msgs::msg::EnvironmentData env_data; + + auto a = to_json(FaultHandlers::build_sovd_fault_response(fault_json(first), env_json(env_data), "/apps/lidar")); + auto b = to_json(FaultHandlers::build_sovd_fault_response(fault_json(second), env_json(env_data), "/apps/motor")); + + EXPECT_EQ(a["item"]["code"], b["item"]["code"]); + EXPECT_EQ(a["x-medkit"]["owner"], "/perception/lidar"); + EXPECT_EQ(b["x-medkit"]["owner"], "/control/motor"); } // @verifies REQ_INTEROP_013 @@ -555,9 +587,10 @@ TEST(FaultInSourceScopeTest, BareEntityIdScopeRejectsOtherEntityFault) { } TEST(FaultInSourceScopeTest, BareEntityIdScopeRejectsMixedSourceFault) { - // Cross-entity clear/disclosure stays blocked under the all-sources rule: a - // fault co-reported by "process" and another entity is out of scope for - // "process" alone. + // Cross-entity clear/disclosure stays blocked under the all-sources rule. A + // record names one source, but one relayed from a peer or read from an older + // store may carry two, and one that names "process" and another entity is + // out of scope for "process" alone. EXPECT_FALSE(FaultHandlers::fault_in_source_scope(make_fault({"process", "s7_status"}), {"process"})); } @@ -572,6 +605,222 @@ TEST(FaultInSourceScopeTest, BareEntityIdScopeRejectsPrefixCollision) { // This pins the shape build_sovd_fault_response produces for the PLC case: // reporting_sources=[bare id] + a freeze-frame snapshot under a plugin entity // path. +// ============================================================================= +// select_scoped_fault - which record a per-record route acts on +// +// A fault code is half of a record's identity. These pin what the per-entity +// GET and DELETE do when the entity's own scope holds none, one, or several +// records of the code the URL named. +// ============================================================================= +namespace { + +json record(const std::string & code, const std::string & owner) { + return json{{"fault_code", code}, {"status", "CONFIRMED"}, {"reporting_sources", json::array({owner})}}; +} + +std::vector in_scope(const json & faults, const std::string & code, + const std::set & scope) { + return ros2_medkit_gateway::faults::records_of_code_in_scope(faults, code, scope); +} + +} // namespace + +// @verifies REQ_INTEROP_013 +TEST(SelectScopedFaultTest, OneRecordInScopeIsTheRecordAndNamesItsOwner) { + const json faults = + json::array({record("SHARED_CODE", "app_a"), record("SHARED_CODE", "app_b"), record("OTHER_CODE", "app_a")}); + + auto picked = + FaultHandlers::select_scoped_fault(in_scope(faults, "SHARED_CODE", {"app_a"}), "SHARED_CODE", "app_id", "app_a"); + + ASSERT_TRUE(picked.has_value()) << picked.error().message; + EXPECT_EQ(picked->owner, "app_a"); + EXPECT_EQ(picked->fault["fault_code"], "SHARED_CODE"); +} + +// @verifies REQ_INTEROP_013 +TEST(SelectScopedFaultTest, NoRecordInScopeIs404) { + const json faults = json::array({record("SHARED_CODE", "app_b")}); + + auto picked = + FaultHandlers::select_scoped_fault(in_scope(faults, "SHARED_CODE", {"app_a"}), "SHARED_CODE", "app_id", "app_a"); + + ASSERT_FALSE(picked.has_value()); + EXPECT_EQ(picked.error().http_status, 404); + EXPECT_EQ(picked.error().code, ros2_medkit_gateway::ERR_RESOURCE_NOT_FOUND); +} + +// Two of this component's own apps report the code, so the URL names two +// records. Acting on either would clear or disclose a record the caller did +// not ask for and could not tell apart in the response. +// @verifies REQ_INTEROP_015 +TEST(SelectScopedFaultTest, SeveralRecordsInScopeIs409NamingEveryOwner) { + const json faults = json::array({record("SHARED_CODE", "app_b"), record("SHARED_CODE", "app_a")}); + + auto picked = FaultHandlers::select_scoped_fault(in_scope(faults, "SHARED_CODE", {"app_a", "app_b"}), "SHARED_CODE", + "component_id", "host"); + + ASSERT_FALSE(picked.has_value()); + EXPECT_EQ(picked.error().http_status, 409); + EXPECT_EQ(picked.error().code, ros2_medkit_gateway::ERR_AMBIGUOUS_FAULT); + ASSERT_TRUE(picked.error().params.contains("owners")); + // Ordered by owner, so the answer does not depend on the store's listing order. + EXPECT_EQ(picked.error().params["owners"], json::array({"app_a", "app_b"})); + EXPECT_EQ(picked.error().params["fault_code"], "SHARED_CODE"); + EXPECT_EQ(picked.error().params["component_id"], "host"); +} + +// The records of ONE code only: a second code the entity also owns is a +// different resource and must not make its sibling ambiguous. +TEST(RecordsOfCodeInScopeTest, OtherCodesAreNotCollected) { + const json faults = json::array({record("SHARED_CODE", "app_a"), record("OTHER_CODE", "app_b")}); + + const auto records = in_scope(faults, "SHARED_CODE", {"app_a", "app_b"}); + + ASSERT_EQ(records.size(), 1u); + EXPECT_EQ(records[0].owner, "app_a"); +} + +// The rosbag download reads the same records to find the owners it asks +// whether they hold the recording, and a code addressing several of them is not +// an ambiguity there: every owner in scope is asked. Reading the code unscoped +// instead answered "ambiguous" and the entity was refused a recording it owns. +TEST(RecordsOfCodeInScopeTest, EveryOwnerInScopeIsARecordingDownloadCandidate) { + const json faults = json::array({record("SHARED_CODE", "app_a"), record("SHARED_CODE", "app_b")}); + + EXPECT_FALSE(in_scope(faults, "SHARED_CODE", {"app_a"}).empty()) << "app_a owns a record of this code"; + EXPECT_EQ(in_scope(faults, "SHARED_CODE", {"app_a", "app_b"}).size(), 2u) + << "a component hosting both owners asks both of them"; + EXPECT_TRUE(in_scope(faults, "SHARED_CODE", {"unrelated_app"}).empty()) + << "an entity owning no record of this code has no owner to ask"; +} + +// The rosbag download resolves the SAME way the fault routes do when the URL +// carries a bare fault code: exactly one record in scope names a recording, and +// several name none of them. Serving the lowest-sorting owner's bag would hand a +// caller another owner's bytes under a URL that never said whose they were. The +// recording-id form addresses the bag directly and is unaffected. +TEST(RecordsOfCodeInScopeTest, SeveralOwnersMakeABareCodeUrlAmbiguous) { + const json faults = json::array({record("SHARED_CODE", "app_b"), record("SHARED_CODE", "app_a")}); + + const auto both = in_scope(faults, "SHARED_CODE", {"app_a", "app_b"}); + ASSERT_EQ(both.size(), 2u) << "a component hosting both owners resolves two records"; + EXPECT_EQ(both[0].owner, "app_a"); + EXPECT_EQ(both[1].owner, "app_b"); + + // One owner in scope is still one record, so that URL keeps working. + const auto single = in_scope(faults, "SHARED_CODE", {"app_a"}); + ASSERT_EQ(single.size(), 1u); + EXPECT_EQ(single[0].owner, "app_a"); +} + +// A muted_faults entry names one record by its code AND its owner. An entry for +// another owner of the same code, or for the same owner under another code, +// hides nothing here, so the one record the entity's list shows is the only +// candidate and the code is not ambiguous. +TEST(AddressableRecordsTest, AMutedEntryHidesOnlyTheRecordItNames) { + const json listing{ + {"faults", json::array({record("SHARED_CODE", "tank"), record("SHARED_CODE", "pump")})}, + {"muted_faults", json::array({json{{"fault_code", "SHARED_CODE"}, {"source_id", "pump"}}, + json{{"fault_code", "OTHER_CODE"}, {"source_id", "tank"}}})}, + }; + + const auto records = ros2_medkit_gateway::faults::addressable_records(listing, "SHARED_CODE", {"tank", "pump"}); + + ASSERT_EQ(records.size(), 1u) << "only pump's record of this code is muted"; + EXPECT_EQ(records[0].owner, "tank"); +} + +// With no shown record of the code the muted ones are the candidates, every one +// of them, so two of them still make the code ambiguous rather than picking one. +TEST(AddressableRecordsTest, MutedRecordsAnswerOnlyWhenNoneIsShown) { + const json listing{ + {"faults", json::array({record("SHARED_CODE", "tank"), record("SHARED_CODE", "pump")})}, + {"muted_faults", json::array({json{{"fault_code", "SHARED_CODE"}, {"source_id", "pump"}}, + json{{"fault_code", "SHARED_CODE"}, {"source_id", "tank"}}})}, + }; + + const auto records = ros2_medkit_gateway::faults::addressable_records(listing, "SHARED_CODE", {"tank", "pump"}); + + ASSERT_EQ(records.size(), 2u); + EXPECT_EQ(records[0].owner, "pump"); + EXPECT_EQ(records[1].owner, "tank"); +} + +namespace { + +json record_in(const std::string & status, const std::string & code, const std::string & owner) { + json r = record(code, owner); + r["status"] = status; + return r; +} + +} // namespace + +// The entity's default fault list shows PREFAILED and CONFIRMED records. A +// record the list only shows under status=cleared or status=healed (CLEARED, +// HEALED, PREPASSED) yields to one it shows by default, so once one of two +// sources has cleared its record the code names the other one again instead of +// answering 409 for good. +TEST(AddressableRecordsTest, ARecordTheListShowsOutranksOneItShowsOnlyOnRequest) { + for (const std::string inactive : {"CLEARED", "HEALED", "PREPASSED"}) { + for (const std::string active : {"CONFIRMED", "PREFAILED"}) { + const json listing{ + {"faults", + json::array({record_in(inactive, "SHARED_CODE", "tank"), record_in(active, "SHARED_CODE", "pump")})}, + }; + + const auto records = ros2_medkit_gateway::faults::addressable_records(listing, "SHARED_CODE", {"tank", "pump"}); + + ASSERT_EQ(records.size(), 1u) << inactive << " beside " << active; + EXPECT_EQ(records[0].owner, "pump") << inactive << " beside " << active; + } + } +} + +// A muted record is hidden from the list only because a root cause muted it. +// It still outranks a record the list hides for its status. +TEST(AddressableRecordsTest, AMutedActiveRecordOutranksAClearedOne) { + const json listing{ + {"faults", + json::array({record_in("CONFIRMED", "SHARED_CODE", "tank"), record_in("CLEARED", "SHARED_CODE", "pump")})}, + {"muted_faults", json::array({json{{"fault_code", "SHARED_CODE"}, {"source_id", "tank"}}})}, + }; + + const auto records = ros2_medkit_gateway::faults::addressable_records(listing, "SHARED_CODE", {"tank", "pump"}); + + ASSERT_EQ(records.size(), 1u); + EXPECT_EQ(records[0].owner, "tank"); +} + +// Nothing the list shows and nothing muted: the cleared and healed records are +// the candidates, all of them. One alone is addressable, two make the code +// ambiguous. A muted record the list would not show anyway (here HEALED) is in +// that last tier too, because its status hides it from the list whether it is +// muted or not. +TEST(AddressableRecordsTest, ClearedAndHealedRecordsAnswerOnlyWhenNothingElseDoes) { + const json one{{"faults", json::array({record_in("CLEARED", "SHARED_CODE", "tank")})}}; + const auto single = ros2_medkit_gateway::faults::addressable_records(one, "SHARED_CODE", {"tank", "pump"}); + ASSERT_EQ(single.size(), 1u) << "one cleared record alone is still the record the code names"; + EXPECT_EQ(single[0].owner, "tank"); + + const json two{ + {"faults", + json::array({record_in("CLEARED", "SHARED_CODE", "tank"), record_in("HEALED", "SHARED_CODE", "pump")})}, + {"muted_faults", json::array({json{{"fault_code", "SHARED_CODE"}, {"source_id", "pump"}}})}, + }; + const auto both = ros2_medkit_gateway::faults::addressable_records(two, "SHARED_CODE", {"tank", "pump"}); + ASSERT_EQ(both.size(), 2u) << "two inactive records name neither"; + EXPECT_EQ(both[0].owner, "pump"); + EXPECT_EQ(both[1].owner, "tank"); +} + +TEST(RecordsOfCodeInScopeTest, RecordOwnerIsTheSingleReportingSource) { + EXPECT_EQ(ros2_medkit_gateway::faults::record_owner(record("C", "app_a")), "app_a"); + EXPECT_EQ(ros2_medkit_gateway::faults::record_owner(json{{"fault_code", "C"}}), ""); + EXPECT_EQ(ros2_medkit_gateway::faults::record_owner(json{{"reporting_sources", json::array()}}), ""); +} + TEST_F(FaultHandlersTest, BuildSovdFaultResponseExternalEntityFreezeFrame) { ros2_medkit_msgs::msg::Fault fault; fault.fault_code = "PLC_LEVEL_HIGH"; @@ -616,8 +865,8 @@ TEST(FaultListItemSchema, FaultToJsonConformsAndRoundTrips) { fault.description = "Brake pressure below threshold"; fault.occurrence_count = 3; fault.status = "active"; - fault.reporting_sources = {"brake_ecu", "abs_node"}; - fault.last_passed.sec = 1200; // absent-when-zero covered separately below + fault.reporting_sources = {"brake_ecu"}; // a record carries its one owner + fault.last_passed.sec = 1200; // absent-when-zero covered separately below const json wire = conversions::fault_to_json(fault); @@ -628,6 +877,34 @@ TEST(FaultListItemSchema, FaultToJsonConformsAndRoundTrips) { EXPECT_EQ(dto::JsonWriter::write(parsed.value()), wire); } +TEST(FaultListItemSchema, FlatItemNamesTheOwningSource) { + // The flat list item addresses its own record: source_id is the owner, and it + // is the single entry of reporting_sources. A client filtering or clearing + // from a list never has to reach into the array to find out whose record it + // is holding. + ros2_medkit_msgs::msg::Fault fault; + fault.fault_code = "SHARED_CODE"; + fault.status = "CONFIRMED"; + fault.reporting_sources = {"app_a"}; + + const json wire = conversions::fault_to_json(fault); + + ASSERT_TRUE(wire.contains("source_id")) << "flat fault item must name the record owner"; + EXPECT_EQ(wire["source_id"], "app_a"); + ASSERT_EQ(wire["reporting_sources"].size(), 1u); + EXPECT_EQ(wire["reporting_sources"][0], "app_a"); +} + +TEST(FaultListItemSchema, FlatItemOmitsSourceIdWithoutAnOwner) { + // A record always has an owner, but the conversion is fed straight from the + // wire and must not invent one: no reporting source, no source_id key. + ros2_medkit_msgs::msg::Fault fault; + fault.fault_code = "ORPHANED"; + fault.status = "CONFIRMED"; + + EXPECT_FALSE(conversions::fault_to_json(fault).contains("source_id")); +} + TEST(FaultListItemSchema, LastPassedOmittedWhenNeverPassed) { // last_passed carries the last PASSED instant; zero means the fault never // reported PASSED, and the wire says that by omitting the key entirely. diff --git a/src/ros2_medkit_gateway/test/test_fault_handlers_plugin_clear.cpp b/src/ros2_medkit_gateway/test/test_fault_handlers_plugin_clear.cpp new file mode 100644 index 000000000..a4c575a6e --- /dev/null +++ b/src/ros2_medkit_gateway/test/test_fault_handlers_plugin_clear.cpp @@ -0,0 +1,1190 @@ +// Copyright 2026 bburda +// +// Licensed under the Apache License, Version 2.0 (the "License"); +// you may not use this file except in compliance with the License. +// You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, software +// distributed under the License is distributed on an "AS IS" BASIS, +// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +// See the License for the specific language governing permissions and +// limitations under the License. + +// The per-record routes, driven through the handlers with a real GatewayNode: +// the fault detail and clear (native and plugin provider) and the recording +// download by fault code. +// +// The other fault-handler suites test pure helpers. These need the handler +// itself, because what they pin is which record the handler resolves and which +// owner it hands on: a record is (fault_code, owner), the owner is what the +// gateway resolved in the entity's fault scope, and for a COMPONENT that owner +// is the hosted APP, not the component in the URL. Handing the entity id down +// instead addresses a record nobody owns, the fault manager declines it, and +// the route still answers 2xx with the record untouched. +// +// The fault manager is a stub service on a second node, the same shape +// test_fault_manager.cpp uses. It answers ListFaults the way the real one +// does: a muted record is left out unless the listing asks for muted records, +// and then it is also named in `muted_faults`. The plugin is a mock +// FaultProvider added to the node's own PluginManager. + +#include + +#include +#include +#include +#include + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "ros2_medkit_gateway/core/discovery/models/app.hpp" +#include "ros2_medkit_gateway/core/discovery/models/component.hpp" +#include "ros2_medkit_gateway/core/http/error_codes.hpp" +#include "ros2_medkit_gateway/core/http/handlers/bulkdata_handlers.hpp" +#include "ros2_medkit_gateway/core/models/thread_safe_entity_cache.hpp" +#include "ros2_medkit_gateway/core/plugins/plugin_manager.hpp" +#include "ros2_medkit_gateway/core/providers/fault_provider.hpp" +#include "ros2_medkit_gateway/gateway_node.hpp" +#include "ros2_medkit_gateway/http/handlers/fault_handlers.hpp" +#include "ros2_medkit_gateway/http/handlers/handler_context.hpp" +#include "ros2_medkit_msgs/msg/muted_fault_info.hpp" +#include "ros2_medkit_msgs/srv/clear_fault.hpp" +#include "ros2_medkit_msgs/srv/get_fault.hpp" +#include "ros2_medkit_msgs/srv/get_rosbag.hpp" +#include "ros2_medkit_msgs/srv/list_faults.hpp" +#include "ros2_medkit_msgs/srv/list_rosbags.hpp" + +using json = nlohmann::json; +using namespace std::chrono_literals; +using ros2_medkit_gateway::App; +using ros2_medkit_gateway::AuthConfig; +using ros2_medkit_gateway::Component; +using ros2_medkit_gateway::CorsConfig; +using ros2_medkit_gateway::FaultProvider; +using ros2_medkit_gateway::FaultProviderErrorInfo; +using ros2_medkit_gateway::GatewayNode; +using ros2_medkit_gateway::GatewayPlugin; +using ros2_medkit_gateway::PluginManager; +using ros2_medkit_gateway::ThreadSafeEntityCache; +using ros2_medkit_gateway::TlsConfig; +using ros2_medkit_gateway::dto::FaultClearResult; +using ros2_medkit_gateway::dto::FaultDetailResult; +using ros2_medkit_gateway::dto::FaultListResult; +using ros2_medkit_gateway::handlers::BulkDataHandlers; +using ros2_medkit_gateway::handlers::FaultHandlers; +using ros2_medkit_gateway::handlers::HandlerContext; +using ros2_medkit_msgs::srv::GetFault; +using ros2_medkit_msgs::srv::GetRosbag; +using ros2_medkit_msgs::srv::ListFaults; +using ros2_medkit_msgs::srv::ListRosbags; + +namespace { + +constexpr const char * kCode = "SHARED_CODE"; +constexpr const char * kComponent = "device_hub"; +constexpr const char * kHostedApp = "tank"; +constexpr const char * kOtherApp = "pump"; + +/// Records the (entity_id, fault_code, owner) triple of every clear it is asked +/// for, which is the whole point of the fixture. +class RecordingFaultPlugin : public GatewayPlugin, public FaultProvider { + public: + struct ClearCall { + std::string entity_id; + std::string fault_code; + std::string owner; + }; + + std::string name() const override { + return "recording_fault_plugin"; + } + void configure(const json & /*config*/) override { + } + void shutdown() override { + } + + tl::expected list_faults(const std::string & /*entity_id*/) override { + return FaultListResult{json{{"items", listed_items_}}}; + } + tl::expected get_fault(const std::string & /*entity_id*/, + const std::string & code) override { + return FaultDetailResult{json{{"code", code}}}; + } + // The gateway clears through clear_fault_record, so a call landing here means + // it went back to the code-only clear and dropped the owner it resolved. + tl::expected clear_fault(const std::string & /*entity_id*/, + const std::string & code) override { + ADD_FAILURE() << "the gateway cleared '" << code << "' by code alone instead of calling clear_fault_record"; + return FaultClearResult{json{{"code", code}, {"cleared", true}}}; + } + tl::expected + clear_fault_record(const std::string & entity_id, const std::string & code, const std::string & owner) override { + if (fail_if_called) { + ADD_FAILURE() << "the provider was asked to clear '" << code << "' with owner '" << owner + << "', which the gateway never resolved"; + } + clears.push_back(ClearCall{entity_id, code, owner}); + return FaultClearResult{json{{"code", code}, {"cleared", true}}}; + } + + std::vector clears; + json listed_items_ = json::array(); + /// Set by a test in which any clear reaching the provider is itself the defect. + bool fail_if_called = false; +}; + +/// A provider written against the two-argument clear contract: it overrides +/// clear_fault(entity_id, fault_code) and nothing else of the clear. It has to +/// keep compiling against the current headers and keep receiving the clear. +class TwoArgumentClearPlugin : public GatewayPlugin, public FaultProvider { + public: + struct ClearCall { + std::string entity_id; + std::string fault_code; + }; + + std::string name() const override { + return "two_argument_clear_plugin"; + } + void configure(const json & /*config*/) override { + } + void shutdown() override { + } + + tl::expected list_faults(const std::string & /*entity_id*/) override { + return FaultListResult{json{{"items", json::array()}}}; + } + tl::expected get_fault(const std::string & /*entity_id*/, + const std::string & code) override { + return FaultDetailResult{json{{"code", code}}}; + } + tl::expected clear_fault(const std::string & entity_id, + const std::string & code) override { + clears.push_back(ClearCall{entity_id, code}); + return FaultClearResult{json{{"code", code}, {"cleared", true}}}; + } + + std::vector clears; +}; + +/// A TypedRequest carrying the two positional captures the fault routes read. +/// The smatch holds iterators into `path`, so both outlive the request. +class RoutedRequest { + public: + RoutedRequest(std::string path, const std::string & pattern, const std::string & method = "DELETE") + : path_(std::move(path)) { + std::regex re(pattern); + matched_ = std::regex_match(path_, raw_.matches, re); + raw_.path = path_; + raw_.method = method; + } + + bool matched() const { + return matched_; + } + const httplib::Request & raw() const { + return raw_; + } + + private: + std::string path_; + httplib::Request raw_; + bool matched_ = false; +}; + +} // namespace + +class PluginClearOwnerTest : public ::testing::Test { + protected: + /// A recording the stub fault manager holds for one record. + struct Recording { + std::string id; + std::string fault_code; + std::string owner; + std::string path; + }; + + static void SetUpTestSuite() { + rclcpp::init(0, nullptr); + } + static void TearDownTestSuite() { + rclcpp::shutdown(); + } + + void SetUp() override { + const int test_id = test_counter_++; + ns_ = "pco" + std::to_string(test_id); + bag_dir_ = std::filesystem::temp_directory_path() / ("pco_bags_" + std::to_string(::getpid()) + "_" + ns_); + std::filesystem::create_directories(bag_dir_); + + auto options = rclcpp::NodeOptions{}.automatically_declare_parameters_from_overrides(false).parameter_overrides({ + {"server.port", 0}, + {"fault_manager.namespace", ns_}, + {"fault_manager.service_timeout_sec", 5.0}, + }); + node_ = std::make_shared(options); + // The injected cache is the single source of truth for these tests. The + // graph-event refresh would reconcile it back to the live ROS graph. + node_->stop_discovery_refresh_for_testing(); + + store_ = std::make_shared("pco_store_" + std::to_string(test_id)); + const std::string base = "/" + ns_ + "/fault_manager/"; + list_srv_ = store_->create_service( + base + "list_faults", + [this](const std::shared_ptr & req, const std::shared_ptr & res) { + std::lock_guard lock(store_mutex_); + last_include_muted_.store(req->include_muted); + res->faults = stored_faults_; + res->muted_count = static_cast(muted_faults_.size()); + // The real fault manager's contract: a muted record is one the + // correlation engine hid from the default listing, so it is served + // only to a listing that asks, and then named in muted_faults by its + // code and owner. + if (req->include_muted) { + for (const auto & muted : muted_faults_) { + res->faults.push_back(muted); + ros2_medkit_msgs::msg::MutedFaultInfo info; + info.fault_code = muted.fault_code; + info.source_id = muted.reporting_sources.front(); + info.root_cause_code = "ROOT_CAUSE"; + info.rule_id = "rule"; + res->muted_faults.push_back(info); + } + } + }); + get_srv_ = + store_->create_service(base + "get_fault", [this](const std::shared_ptr & req, + const std::shared_ptr & res) { + std::lock_guard lock(store_mutex_); + get_requests_.emplace_back(req->fault_code, req->source_id); + for (const auto * list : {&stored_faults_, &muted_faults_}) { + for (const auto & fault : *list) { + if (fault.fault_code == req->fault_code && fault.reporting_sources.front() == req->source_id) { + res->success = true; + res->fault = fault; + return; + } + } + } + res->success = false; + res->error_message = "Fault not found: " + req->fault_code; + }); + clear_srv_ = store_->create_service( + base + "clear_fault", [this](const std::shared_ptr & req, + const std::shared_ptr & res) { + { + std::lock_guard lock(cleared_mutex_); + cleared_.push_back({req->fault_code, req->source_id}); + } + // The record the request names moves to CLEARED, so a test can read + // the outcome off the store rather than off the request log alone. + std::lock_guard lock(store_mutex_); + for (auto * list : {&stored_faults_, &muted_faults_}) { + for (auto & fault : *list) { + if (fault.fault_code == req->fault_code && fault.reporting_sources.front() == req->source_id) { + fault.status = ros2_medkit_msgs::msg::Fault::STATUS_CLEARED; + } + } + } + res->success = true; + res->message = "cleared"; + }); + // The fault manager's lookup order: the recording id first, then the fault + // code scoped to the owner the gateway sent. An unscoped code with more than + // one owner holding recordings answers nothing. + rosbag_srv_ = store_->create_service( + base + "get_rosbag", + [this](const std::shared_ptr & req, const std::shared_ptr & res) { + std::lock_guard lock(store_mutex_); + rosbag_requests_.push_back(req->source_id); + const Recording * found = nullptr; + for (const auto & rec : recordings_) { + if (rec.id == req->recording_id) { + found = &rec; + } + } + if (found == nullptr) { + std::vector by_code; + for (const auto & rec : recordings_) { + if (rec.fault_code == req->fault_code && (req->source_id.empty() || rec.owner == req->source_id)) { + by_code.push_back(&rec); + } + } + if (by_code.size() == 1) { + found = by_code.front(); + } + } + if (found == nullptr) { + res->success = false; + res->error_message = "No recording for " + req->recording_id; + return; + } + res->success = true; + res->file_path = found->path; + res->recording_id = found->id; + res->fault_codes = {found->fault_code}; + res->format = "mcap"; + res->size_bytes = std::filesystem::file_size(found->path); + }); + list_rosbags_srv_ = store_->create_service( + base + "list_rosbags", + [this](const std::shared_ptr & req, const std::shared_ptr & res) { + std::lock_guard lock(store_mutex_); + res->success = true; + for (const auto & rec : recordings_) { + if (rec.owner != req->entity_fqn) { + continue; + } + res->fault_codes.push_back(rec.fault_code); + res->recording_ids.push_back(rec.id); + res->file_paths.push_back(rec.path); + res->formats.push_back("mcap"); + res->durations_sec.push_back(1.0); + res->sizes_bytes.push_back(std::filesystem::file_size(rec.path)); + res->created_at_ns.push_back(1'700'000'000'000'000'000); + } + }); + executor_ = std::make_unique(); + executor_->add_node(store_); + spin_ = std::thread([this] { + executor_->spin(); + }); + + ctx_ = std::make_unique(node_.get(), cors_, auth_, tls_, nullptr); + handlers_ = std::make_unique(*ctx_); + bulk_handlers_ = std::make_unique(*ctx_); + } + + void TearDown() override { + if (spin_.joinable()) { + executor_->cancel(); + spin_.join(); + } + executor_.reset(); + store_.reset(); + bulk_handlers_.reset(); + handlers_.reset(); + ctx_.reset(); + node_.reset(); + std::error_code ec; + std::filesystem::remove_all(bag_dir_, ec); + } + + static ros2_medkit_msgs::msg::Fault make_record(const std::string & code, const std::string & owner) { + ros2_medkit_msgs::msg::Fault fault; + fault.fault_code = code; + fault.status = ros2_medkit_msgs::msg::Fault::STATUS_CONFIRMED; + fault.severity = ros2_medkit_msgs::msg::Fault::SEVERITY_ERROR; + fault.reporting_sources = {owner}; + return fault; + } + + /// One stored record, owned by `owner`. + void store_record(const std::string & code, const std::string & owner) { + std::lock_guard lock(store_mutex_); + stored_faults_.push_back(make_record(code, owner)); + } + + /// One muted record, owned by `owner`. Served only to a listing that asks. + void store_muted_record(const std::string & code, const std::string & owner) { + std::lock_guard lock(store_mutex_); + muted_faults_.push_back(make_record(code, owner)); + } + + /// Move the stored record (code, owner) to `status`, muted or not. + void set_status(const std::string & code, const std::string & owner, const std::string & status) { + std::lock_guard lock(store_mutex_); + for (auto * list : {&stored_faults_, &muted_faults_}) { + for (auto & fault : *list) { + if (fault.fault_code == code && fault.reporting_sources.front() == owner) { + fault.status = status; + } + } + } + } + + /// The HTTP status a recording download answers: 200 when it serves bytes, + /// otherwise the status of the error it returns. + int download_status(const std::string & path, const std::string & pattern) { + RoutedRequest req(path, pattern, "GET"); + EXPECT_TRUE(req.matched()) << path; + auto result = bulk_handlers_->download(ros2_medkit_gateway::http::TypedRequest(req.raw())); + return result.has_value() ? 200 : result.error().http_status; + } + + static std::string app_bag_path(const std::string & app, const std::string & id) { + return std::string("/api/v1/apps/") + app + "/bulk-data/rosbags/" + id; + } + static std::string component_bag_path(const std::string & id) { + return std::string("/api/v1/components/") + kComponent + "/bulk-data/rosbags/" + id; + } + + /// A recording of the record (code, owner), with real bytes behind it. + void store_recording(const std::string & id, const std::string & code, const std::string & owner) { + const auto path = bag_dir_ / (id + ".mcap"); + { + std::ofstream out(path, std::ios::binary); + out << "bag of " << owner; + } + std::lock_guard lock(store_mutex_); + recordings_.push_back(Recording{id, code, owner, path.string()}); + } + + /// A component hosting two external apps, all three owned by the plugin, so + /// the fault routes take the plugin branch on every one of them. + RecordingFaultPlugin * seed_topology() { + return seed_plugin_topology(); + } + + /// The same topology owned by a plugin of type P. + template + P * seed_plugin_topology() { + seed_entities(); + auto plugin = std::make_unique

(); + auto * raw = plugin.get(); + const std::string plugin_name = raw->name(); + auto * pmgr = node_->get_plugin_manager(); + pmgr->add_plugin(std::move(plugin)); + pmgr->register_entity_ownership(plugin_name, {kComponent, kHostedApp, kOtherApp}); + return raw; + } + + /// The same component and apps, owned by NOBODY, so the fault routes take the + /// native fault-manager path and resolve through resolve_scoped_fault. + void seed_native_topology() { + seed_entities(); + } + + bool wait_for_store(std::chrono::milliseconds timeout = 5s) { + auto probe = store_->create_client("/" + ns_ + "/fault_manager/list_faults"); + const auto deadline = std::chrono::steady_clock::now() + timeout; + while (std::chrono::steady_clock::now() < deadline) { + if (probe->service_is_ready()) { + return true; + } + std::this_thread::sleep_for(20ms); + } + return false; + } + + std::vector> cleared_snapshot() { + std::lock_guard lock(cleared_mutex_); + return cleared_; + } + + /// The stored status of the record (code, owner), muted or not. + std::string status_of(const std::string & code, const std::string & owner) { + std::lock_guard lock(store_mutex_); + for (const auto * list : {&stored_faults_, &muted_faults_}) { + for (const auto & fault : *list) { + if (fault.fault_code == code && fault.reporting_sources.front() == owner) { + return fault.status; + } + } + } + return ""; + } + + std::vector> get_requests_snapshot() { + std::lock_guard lock(store_mutex_); + return get_requests_; + } + + std::vector rosbag_requests_snapshot() { + std::lock_guard lock(store_mutex_); + return rosbag_requests_; + } + + static std::string component_fault_path() { + return std::string("/api/v1/components/") + kComponent + "/faults/" + kCode; + } + static constexpr const char * kComponentFaultPattern = R"(/api/v1/components/([^/]+)/faults/([^/]+))"; + static constexpr const char * kComponentBagPattern = R"(/api/v1/components/([^/]+)/bulk-data/([^/]+)/([^/]+))"; + static constexpr const char * kAppBagPattern = R"(/api/v1/apps/([^/]+)/bulk-data/([^/]+)/([^/]+))"; + + static inline int test_counter_ = 0; + CorsConfig cors_{}; + AuthConfig auth_{}; + TlsConfig tls_{}; + std::string ns_; + std::filesystem::path bag_dir_; + std::shared_ptr node_; + std::shared_ptr store_; + rclcpp::Service::SharedPtr list_srv_; + rclcpp::Service::SharedPtr get_srv_; + rclcpp::Service::SharedPtr clear_srv_; + rclcpp::Service::SharedPtr rosbag_srv_; + rclcpp::Service::SharedPtr list_rosbags_srv_; + std::unique_ptr executor_; + std::thread spin_; + + std::mutex store_mutex_; + std::vector stored_faults_; + std::vector muted_faults_; + std::vector recordings_; + std::vector> get_requests_; + std::vector rosbag_requests_; + std::atomic last_include_muted_{false}; + + std::mutex cleared_mutex_; + std::vector> cleared_; + std::unique_ptr ctx_; + std::unique_ptr handlers_; + std::unique_ptr bulk_handlers_; + + private: + void seed_entities() { + App a; + a.id = kHostedApp; + a.component_id = kComponent; + a.external = true; + App b; + b.id = kOtherApp; + b.component_id = kComponent; + b.external = true; + Component c; + c.id = kComponent; + c.external = true; + auto & cache = const_cast(node_->get_thread_safe_cache()); + cache.update_all({}, {c}, {a, b}, {}); + } +}; + +// A component owns the records its hosted apps reported. The clear the plugin +// receives has to name that app, because that is the record the gateway +// resolved. Handing it the component id addresses a record no source owns. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, ComponentClearHandsTheHostedAppsOwnerToThePlugin) { + auto * plugin = seed_topology(); + store_record(kCode, kHostedApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/components/") + kComponent + "/faults/" + kCode, + R"(/api/v1/components/([^/]+)/faults/([^/]+))"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << result.error().message; + ASSERT_EQ(plugin->clears.size(), 1u); + EXPECT_EQ(plugin->clears[0].entity_id, kComponent); + EXPECT_EQ(plugin->clears[0].fault_code, kCode); + EXPECT_EQ(plugin->clears[0].owner, kHostedApp) << "the clear must name the record's owner, not the entity in the URL"; +} + +// The app route resolves to its own record, so owner and entity coincide there. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, AppClearHandsItsOwnOwnerToThePlugin) { + auto * plugin = seed_topology(); + store_record(kCode, kHostedApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/apps/") + kHostedApp + "/faults/" + kCode, + R"(/api/v1/apps/([^/]+)/faults/([^/]+))"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << result.error().message; + ASSERT_EQ(plugin->clears.size(), 1u); + EXPECT_EQ(plugin->clears[0].owner, kHostedApp); +} + +// Two hosted apps report the code, so the component's URL names neither record +// and the plugin must not be asked to clear anything. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, AmbiguousComponentClearReachesNoProvider) { + auto * plugin = seed_topology(); + store_record(kCode, kHostedApp); + store_record(kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/components/") + kComponent + "/faults/" + kCode, + R"(/api/v1/components/([^/]+)/faults/([^/]+))"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_FALSE(result.has_value()); + EXPECT_EQ(result.error().http_status, 409); + EXPECT_TRUE(plugin->clears.empty()) << "an ambiguous address must not reach the provider"; +} + +// The bulk clear walks the plugin's own listing, and each item names its own +// record. Sending the entity id for all of them clears one owner's record at +// most and silently misses the rest. +// @verifies REQ_INTEROP_014 +TEST_F(PluginClearOwnerTest, ComponentBulkClearSendsEachRecordsOwnOwner) { + auto * plugin = seed_topology(); + plugin->listed_items_ = json::array({ + json{{"code", kCode}, {"source_id", kHostedApp}}, + json{{"code", kCode}, {"source_id", kOtherApp}}, + }); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/components/") + kComponent + "/faults", + R"(/api/v1/components/([^/]+)/faults)"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_all_faults(typed); + + ASSERT_TRUE(result.has_value()) << result.error().message; + ASSERT_EQ(plugin->clears.size(), 2u); + EXPECT_EQ(plugin->clears[0].owner, kHostedApp); + EXPECT_EQ(plugin->clears[1].owner, kOtherApp) << "each item's own source_id must travel with its clear"; +} + +// The correlation engine mutes a symptom to keep it out of the default listing. +// That is a display decision, not an unaddressing: the record still exists and +// its entity could read and clear it before. resolve_scoped_fault resolving out +// of a listing that excludes muted records 404s it instead. +// +// The topology here is deliberately NOT plugin-owned, so the route takes the +// native fault-manager path and resolve_scoped_fault is the function under test. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, ResolutionKeepsAMutedRecordAddressable) { + seed_native_topology(); + store_muted_record(kCode, kHostedApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/apps/") + kHostedApp + "/faults/" + kCode, + R"(/api/v1/apps/([^/]+)/faults/([^/]+))"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << "a muted record must stay addressable: " << result.error().message; + const auto cleared = cleared_snapshot(); + ASSERT_EQ(cleared.size(), 1u) << "the muted record was never cleared"; + EXPECT_EQ(cleared[0].first, kCode); + EXPECT_EQ(cleared[0].second, kHostedApp) << "the clear must carry the muted record's owner"; +} + +// The same route on an unmuted record, so the muted case above is not simply +// passing because the harness serves everything regardless of the flag. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, ResolutionClearsAnUnmutedRecordWithItsOwner) { + seed_native_topology(); + store_record(kCode, kHostedApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/apps/") + kHostedApp + "/faults/" + kCode, + R"(/api/v1/apps/([^/]+)/faults/([^/]+))"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << result.error().message; + const auto cleared = cleared_snapshot(); + ASSERT_EQ(cleared.size(), 1u); + EXPECT_EQ(cleared[0].second, kHostedApp); +} + +// The plugin clear reads the fault manager to learn which record the code +// names. When that read fails the route cannot tell a record the fault manager +// holds from a plugin-internal one, so it must not guess: calling the provider +// with an empty owner answered 2xx while the record stayed untouched. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, PluginClearAnswers503WhenTheFaultManagerCannotBeRead) { + auto * plugin = seed_topology(); + plugin->fail_if_called = true; + store_record(kCode, kHostedApp); + ASSERT_TRUE(wait_for_store()); + // Take ListFaults away and wait until the graph agrees it is gone, so the + // call below fails on the read and not on a race with the service teardown. + list_srv_.reset(); + auto probe = store_->create_client("/" + ns_ + "/fault_manager/list_faults"); + const auto deadline = std::chrono::steady_clock::now() + 5s; + while (probe->service_is_ready() && std::chrono::steady_clock::now() < deadline) { + std::this_thread::sleep_for(20ms); + } + ASSERT_FALSE(probe->service_is_ready()) << "ListFaults is still discoverable"; + + RoutedRequest req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_FALSE(result.has_value()) << "a clear the gateway could not resolve must not report success"; + EXPECT_EQ(result.error().http_status, 503); + EXPECT_EQ(result.error().code, ros2_medkit_gateway::ERR_SERVICE_UNAVAILABLE); + EXPECT_TRUE(plugin->clears.empty()) << "the provider was called without a resolved owner"; +} + +// The positive control for the case above, on the same fixture and the same +// stub: with ListFaults answering, the same request reaches the provider with +// the owner the gateway resolved. So the silence above is the route refusing, +// not a harness that never reaches the provider. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, PluginClearReachesTheProviderWhenTheFaultManagerAnswers) { + auto * plugin = seed_topology(); + store_record(kCode, kHostedApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << result.error().message; + ASSERT_EQ(plugin->clears.size(), 1u); + EXPECT_EQ(plugin->clears[0].owner, kHostedApp); +} + +// A provider that predates the owner-aware clear overrides only the +// two-argument clear_fault. It must still compile against these headers and +// still receive the clear the gateway resolved, through the default +// clear_fault_record, which hands it the entity and the code as before. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, ATwoArgumentClearProviderStillReceivesTheClear) { + auto * plugin = seed_plugin_topology(); + store_record(kCode, kHostedApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << result.error().message; + ASSERT_TRUE(std::holds_alternative(*result)) << "the plugin's acknowledgement is the answer"; + ASSERT_EQ(plugin->clears.size(), 1u) << "the two-argument clear_fault never received the clear"; + EXPECT_EQ(plugin->clears[0].entity_id, kComponent); + EXPECT_EQ(plugin->clears[0].fault_code, kCode); +} + +// --------------------------------------------------------------------------- +// Which record a code in the URL names when muted records are in scope. +// +// Every per-entity fault list leaves muted records out, so the records it shows +// are the ones a client can read a code off. A code resolves over those first. +// Only when the list shows no record of the code does a muted one answer, so a +// record hidden as a symptom stays reachable without turning a shown record +// into an ambiguous address. Each route that resolves a code gets its own case, +// because each reads the fault manager on its own. +// --------------------------------------------------------------------------- + +// The component's list shows tank's record only, because pump's record of the +// same code is muted. Resolving over both answered 409 naming pump as well, and +// pointed the client at a list that could never show it pump's record. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, NativeComponentClearTakesTheShownRecordOverAMutedOne) { + seed_native_topology(); + store_record(kCode, kHostedApp); + store_muted_record(kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << "the shown record is the one the code names: " << result.error().message; + EXPECT_TRUE(std::holds_alternative(*result)); + const auto cleared = cleared_snapshot(); + ASSERT_EQ(cleared.size(), 1u) << "exactly the shown record is cleared, and pump's muted record is left alone"; + EXPECT_EQ(cleared[0].first, kCode); + EXPECT_EQ(cleared[0].second, kHostedApp); +} + +// The same scenario on a plugin-owned component: the plugin clear resolves the +// owner on its own read of the fault manager, so it needs its own case. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, PluginComponentClearTakesTheShownRecordOverAMutedOne) { + auto * plugin = seed_topology(); + store_record(kCode, kHostedApp); + store_muted_record(kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << "the shown record is the one the code names: " << result.error().message; + ASSERT_EQ(plugin->clears.size(), 1u); + EXPECT_EQ(plugin->clears[0].owner, kHostedApp) << "the plugin must be handed the shown record's owner"; +} + +// A muted record with no shown record of its code is still the record the code +// names, and the detail route serves it with its owner. +// @verifies REQ_INTEROP_013 +TEST_F(PluginClearOwnerTest, AMutedRecordAloneIsServedOnTheDetailRoute) { + seed_native_topology(); + store_muted_record(kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern, "GET"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->get_fault(typed); + + ASSERT_TRUE(result.has_value()) << "a muted record must stay addressable: " << result.error().message; + EXPECT_EQ(result->content["x-medkit"]["owner"], kOtherApp); + const auto reads = get_requests_snapshot(); + ASSERT_EQ(reads.size(), 1u); + EXPECT_EQ(reads[0].second, kOtherApp) << "the enriched read must be addressed to the muted record's owner"; +} + +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, AMutedRecordAloneIsClearedOnTheNativeRoute) { + seed_native_topology(); + store_muted_record(kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << "a muted record must stay addressable: " << result.error().message; + const auto cleared = cleared_snapshot(); + ASSERT_EQ(cleared.size(), 1u); + EXPECT_EQ(cleared[0].second, kOtherApp); +} + +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, AMutedRecordAloneReachesThePluginWithItsOwner) { + auto * plugin = seed_topology(); + store_muted_record(kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_TRUE(result.has_value()) << result.error().message; + ASSERT_EQ(plugin->clears.size(), 1u); + EXPECT_EQ(plugin->clears[0].owner, kOtherApp) << "the plugin must be handed the muted record's owner"; +} + +// A recording URL carrying a fault code resolves that code to a record first. +// A muted record keeps its recordings, so its code still serves its bytes. +// @verifies REQ_INTEROP_072 +TEST_F(PluginClearOwnerTest, AMutedRecordAloneServesItsRecordingByCode) { + seed_native_topology(); + store_muted_record(kCode, kOtherApp); + store_recording("rec_pump", kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/components/") + kComponent + "/bulk-data/rosbags/" + kCode, + kComponentBagPattern, "GET"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = bulk_handlers_->download(typed); + + ASSERT_TRUE(result.has_value()) << "a muted record's recording must stay reachable: " << result.error().message; + EXPECT_EQ(result->filename.value_or(""), "rec_pump.mcap"); + const auto asked = rosbag_requests_snapshot(); + ASSERT_EQ(asked.size(), 1u); + EXPECT_EQ(asked[0], kOtherApp) << "the recording must be asked for under the muted record's owner"; +} + +// The recording route resolves a code the way the fault routes do: tank's +// record is the one the component's list shows, so the code names tank's +// recording and pump's muted record does not make it ambiguous. +// @verifies REQ_INTEROP_072 +TEST_F(PluginClearOwnerTest, ARecordingCodeTakesTheShownRecordOverAMutedOne) { + seed_native_topology(); + store_record(kCode, kHostedApp); + store_muted_record(kCode, kOtherApp); + store_recording("rec_tank", kCode, kHostedApp); + store_recording("rec_pump", kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/components/") + kComponent + "/bulk-data/rosbags/" + kCode, + kComponentBagPattern, "GET"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = bulk_handlers_->download(typed); + + ASSERT_TRUE(result.has_value()) << "the shown record is the one the code names: " << result.error().message; + EXPECT_EQ(result->filename.value_or(""), "rec_tank.mcap"); + const auto asked = rosbag_requests_snapshot(); + ASSERT_EQ(asked.size(), 1u); + EXPECT_EQ(asked[0], kHostedApp); +} + +// Two muted records and no shown one: the code names neither, the route says so +// and names both owners, and it points at where each can be addressed. +// @verifies REQ_INTEROP_015 +TEST_F(PluginClearOwnerTest, TwoMutedRecordsAloneAreAmbiguous) { + seed_native_topology(); + store_muted_record(kCode, kHostedApp); + store_muted_record(kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = handlers_->clear_fault(typed); + + ASSERT_FALSE(result.has_value()); + EXPECT_EQ(result.error().http_status, 409); + EXPECT_EQ(result.error().params["owners"], json::array({kOtherApp, kHostedApp})); + const std::string details = result.error().params.value("details", ""); + EXPECT_NE(details.find("parameters.owners"), std::string::npos) << details; + EXPECT_NE(details.find("/apps/{app_id}/faults/{fault_code}"), std::string::npos) << details; + EXPECT_TRUE(cleared_snapshot().empty()) << "an ambiguous address must clear nothing"; +} + +// @verifies REQ_INTEROP_072 +TEST_F(PluginClearOwnerTest, TwoMutedRecordsAloneMakeTheRecordingCodeAmbiguous) { + seed_native_topology(); + store_muted_record(kCode, kHostedApp); + store_muted_record(kCode, kOtherApp); + store_recording("rec_tank", kCode, kHostedApp); + store_recording("rec_pump", kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(std::string("/api/v1/components/") + kComponent + "/bulk-data/rosbags/" + kCode, + kComponentBagPattern, "GET"); + ASSERT_TRUE(req.matched()); + ros2_medkit_gateway::http::TypedRequest typed(req.raw()); + + auto result = bulk_handlers_->download(typed); + + ASSERT_FALSE(result.has_value()); + EXPECT_EQ(result.error().http_status, 409); + EXPECT_EQ(result.error().params["owners"], json::array({kOtherApp, kHostedApp})); + EXPECT_TRUE(rosbag_requests_snapshot().empty()) << "no recording is looked up for an ambiguous code"; +} + +// The bulk clear walks the entity's fault list, which leaves muted records out, +// and it always has. It clears what the list shows and leaves a muted record +// CONFIRMED. That record is not stranded: its own per-code DELETE clears it. +// @verifies REQ_INTEROP_014 +TEST_F(PluginClearOwnerTest, BulkClearLeavesAMutedRecordToItsPerCodeClear) { + seed_native_topology(); + store_record(kCode, kHostedApp); + store_muted_record(kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest bulk(std::string("/api/v1/components/") + kComponent + "/faults", + R"(/api/v1/components/([^/]+)/faults)"); + ASSERT_TRUE(bulk.matched()); + auto bulk_result = handlers_->clear_all_faults(ros2_medkit_gateway::http::TypedRequest(bulk.raw())); + + ASSERT_TRUE(bulk_result.has_value()) << bulk_result.error().message; + EXPECT_EQ(status_of(kCode, kHostedApp), "CLEARED") << "the bulk clear must clear the record the list shows"; + EXPECT_EQ(status_of(kCode, kOtherApp), "CONFIRMED") << "the bulk clear must leave the muted record alone"; + + RoutedRequest per_code(std::string("/api/v1/apps/") + kOtherApp + "/faults/" + kCode, + R"(/api/v1/apps/([^/]+)/faults/([^/]+))"); + ASSERT_TRUE(per_code.matched()); + auto per_code_result = handlers_->clear_fault(ros2_medkit_gateway::http::TypedRequest(per_code.raw())); + + ASSERT_TRUE(per_code_result.has_value()) << per_code_result.error().message; + EXPECT_EQ(status_of(kCode, kOtherApp), "CLEARED") << "the muted record's own per-code DELETE must clear it"; +} + +// The recording route's 409 has to send the client somewhere that works. It +// names the rosbags listing and the fault detail's bulk_data_uri, and following +// it does work: the listing carries one descriptor per recording, and each +// descriptor id downloads that owner's bytes. +// @verifies REQ_INTEROP_072 +TEST_F(PluginClearOwnerTest, TheAmbiguousRecordingCodeNamesWhereEachRecordingIsListed) { + seed_native_topology(); + store_record(kCode, kHostedApp); + store_record(kCode, kOtherApp); + store_recording("rec_tank", kCode, kHostedApp); + store_recording("rec_pump", kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + const std::string rosbags = std::string("/api/v1/components/") + kComponent + "/bulk-data/rosbags"; + + RoutedRequest by_code(rosbags + "/" + kCode, kComponentBagPattern, "GET"); + ASSERT_TRUE(by_code.matched()); + auto refused = bulk_handlers_->download(ros2_medkit_gateway::http::TypedRequest(by_code.raw())); + ASSERT_FALSE(refused.has_value()); + ASSERT_EQ(refused.error().http_status, 409); + const std::string details = refused.error().params.value("details", ""); + EXPECT_NE(details.find("GET .../bulk-data/rosbags"), std::string::npos) << details; + EXPECT_NE(details.find("descriptor's id"), std::string::npos) << details; + EXPECT_NE(details.find("environment_data.snapshots[].bulk_data_uri"), std::string::npos) << details; + EXPECT_EQ(details.find("x-medkit.recording_id"), std::string::npos) + << "the fault listing carries no recording id, so the message must not send the client there: " << details; + + // Follow the message: list the category, then download each descriptor id. + RoutedRequest listing(rosbags, R"(/api/v1/components/([^/]+)/bulk-data/([^/]+))", "GET"); + ASSERT_TRUE(listing.matched()); + auto listed = bulk_handlers_->list_descriptors(ros2_medkit_gateway::http::TypedRequest(listing.raw())); + ASSERT_TRUE(listed.has_value()) << listed.error().message; + std::vector ids; + for (const auto & descriptor : listed->items) { + ids.push_back(descriptor.id); + } + std::sort(ids.begin(), ids.end()); + ASSERT_EQ(ids, (std::vector{"rec_pump", "rec_tank"})); + + for (const auto & id : ids) { + RoutedRequest by_id((rosbags + "/").append(id), kComponentBagPattern, "GET"); + ASSERT_TRUE(by_id.matched()); + auto served = bulk_handlers_->download(ros2_medkit_gateway::http::TypedRequest(by_id.raw())); + ASSERT_TRUE(served.has_value()) << id << ": " << served.error().message; + EXPECT_EQ(served->filename.value_or(""), id + ".mcap"); + } +} + +// Two sources report one code and one of them clears its record. The component +// now lists one record of the code, and its per-code routes act on that record +// instead of answering 409 for as long as the cleared record is kept. +// @verifies REQ_INTEROP_013 +TEST_F(PluginClearOwnerTest, AClearedRecordLeavesTheComponentRoutesToTheActiveOne) { + seed_native_topology(); + store_record(kCode, kHostedApp); + store_record(kCode, kOtherApp); + store_recording("rec_tank", kCode, kHostedApp); + store_recording("rec_pump", kCode, kOtherApp); + set_status(kCode, kOtherApp, ros2_medkit_msgs::msg::Fault::STATUS_CLEARED); + ASSERT_TRUE(wait_for_store()); + + // Each route is checked on its own, so one that still answers 409 does not + // hide the others. + RoutedRequest get_req(component_fault_path(), kComponentFaultPattern, "GET"); + ASSERT_TRUE(get_req.matched()); + auto got = handlers_->get_fault(ros2_medkit_gateway::http::TypedRequest(get_req.raw())); + EXPECT_TRUE(got.has_value()) << "GET: the active record is the one the code names: " << got.error().message; + if (got.has_value()) { + EXPECT_EQ(got->content["x-medkit"]["owner"], kHostedApp); + } + + RoutedRequest bag_req(component_bag_path(kCode), kComponentBagPattern, "GET"); + ASSERT_TRUE(bag_req.matched()); + auto served = bulk_handlers_->download(ros2_medkit_gateway::http::TypedRequest(bag_req.raw())); + EXPECT_TRUE(served.has_value()) << "bag URL: the active record's recording is the one the code names: " + << served.error().message; + if (served.has_value()) { + EXPECT_EQ(served->filename.value_or(""), "rec_tank.mcap"); + } + + RoutedRequest del_req(component_fault_path(), kComponentFaultPattern); + ASSERT_TRUE(del_req.matched()); + auto cleared = handlers_->clear_fault(ros2_medkit_gateway::http::TypedRequest(del_req.raw())); + EXPECT_TRUE(cleared.has_value()) << "DELETE: the active record is the one the code clears: " + << cleared.error().message; + const auto clears = cleared_snapshot(); + ASSERT_EQ(clears.size(), cleared.has_value() ? 1u : 0u); + if (cleared.has_value()) { + EXPECT_EQ(clears[0].second, kHostedApp) << "the clear must name the active record's owner"; + EXPECT_EQ(status_of(kCode, kHostedApp), "CLEARED"); + } +} + +// With every record of the code cleared, none is preferred, so two of them name +// neither and the route says so rather than picking one. +// @verifies REQ_INTEROP_013 +TEST_F(PluginClearOwnerTest, TwoClearedRecordsAloneAreAmbiguous) { + seed_native_topology(); + store_record(kCode, kHostedApp); + store_record(kCode, kOtherApp); + set_status(kCode, kHostedApp, ros2_medkit_msgs::msg::Fault::STATUS_CLEARED); + set_status(kCode, kOtherApp, ros2_medkit_msgs::msg::Fault::STATUS_CLEARED); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern, "GET"); + ASSERT_TRUE(req.matched()); + auto result = handlers_->get_fault(ros2_medkit_gateway::http::TypedRequest(req.raw())); + + ASSERT_FALSE(result.has_value()); + EXPECT_EQ(result.error().http_status, 409); + EXPECT_EQ(result.error().params["owners"], json::array({kOtherApp, kHostedApp})); + EXPECT_TRUE(get_requests_snapshot().empty()) << "an ambiguous address must read no record"; +} + +// One cleared record and nothing else: the code still names it, and the detail +// route serves it the way it always served an acknowledged fault. +// @verifies REQ_INTEROP_013 +TEST_F(PluginClearOwnerTest, OneClearedRecordAloneIsServedOnTheDetailRoute) { + seed_native_topology(); + store_record(kCode, kHostedApp); + set_status(kCode, kHostedApp, ros2_medkit_msgs::msg::Fault::STATUS_CLEARED); + ASSERT_TRUE(wait_for_store()); + + RoutedRequest req(component_fault_path(), kComponentFaultPattern, "GET"); + ASSERT_TRUE(req.matched()); + auto result = handlers_->get_fault(ros2_medkit_gateway::http::TypedRequest(req.raw())); + + ASSERT_TRUE(result.has_value()) << "a cleared record alone stays addressable: " << result.error().message; + EXPECT_EQ(result->content["x-medkit"]["owner"], kHostedApp); +} + +// A recording belongs to the (fault code, owner) records it is attached to. Two +// owners of one code each have their own recording: each app downloads its own +// by id and gets 404 on the other's, while the component hosting both owns both. +// The 200s are the positive control for the 404s on the same harness. +// @verifies REQ_INTEROP_072 +TEST_F(PluginClearOwnerTest, EachOwnerDownloadsOnlyItsOwnRecordingById) { + seed_native_topology(); + store_record(kCode, kHostedApp); + store_record(kCode, kOtherApp); + store_recording("rec_tank", kCode, kHostedApp); + store_recording("rec_pump", kCode, kOtherApp); + ASSERT_TRUE(wait_for_store()); + + EXPECT_EQ(download_status(app_bag_path(kHostedApp, "rec_tank"), kAppBagPattern), 200); + EXPECT_EQ(download_status(app_bag_path(kHostedApp, "rec_pump"), kAppBagPattern), 404) + << "tank must not download pump's recording of the same code"; + EXPECT_EQ(download_status(app_bag_path(kOtherApp, "rec_pump"), kAppBagPattern), 200); + EXPECT_EQ(download_status(app_bag_path(kOtherApp, "rec_tank"), kAppBagPattern), 404) + << "pump must not download tank's recording of the same code"; + EXPECT_EQ(download_status(component_bag_path("rec_tank"), kComponentBagPattern), 200); + EXPECT_EQ(download_status(component_bag_path("rec_pump"), kComponentBagPattern), 200); +} + +// Clearing a record keeps its owner in the entity's scope, but it does not make +// that owner the owner of another record's recording. tank cleared its record +// and still gets 404 on pump's recording, while its own recording stays served. +// @verifies REQ_INTEROP_072 +TEST_F(PluginClearOwnerTest, AClearedRecordDoesNotOpenAnotherOwnersRecording) { + seed_native_topology(); + store_record(kCode, kHostedApp); + store_record(kCode, kOtherApp); + store_recording("rec_tank", kCode, kHostedApp); + store_recording("rec_pump", kCode, kOtherApp); + set_status(kCode, kHostedApp, ros2_medkit_msgs::msg::Fault::STATUS_CLEARED); + ASSERT_TRUE(wait_for_store()); + + EXPECT_EQ(download_status(app_bag_path(kHostedApp, "rec_pump"), kAppBagPattern), 404) + << "tank's cleared record must not authorize pump's recording"; + EXPECT_EQ(download_status(app_bag_path(kHostedApp, "rec_tank"), kAppBagPattern), 200) + << "a cleared record keeps serving its own recording"; + EXPECT_EQ(download_status(app_bag_path(kOtherApp, "rec_tank"), kAppBagPattern), 404); + EXPECT_EQ(download_status(app_bag_path(kOtherApp, "rec_pump"), kAppBagPattern), 200); +} + +int main(int argc, char ** argv) { + ::testing::InitGoogleTest(&argc, argv); + return RUN_ALL_TESTS(); +} diff --git a/src/ros2_medkit_gateway/test/test_fault_manager.cpp b/src/ros2_medkit_gateway/test/test_fault_manager.cpp index a2125a806..642f906ca 100644 --- a/src/ros2_medkit_gateway/test/test_fault_manager.cpp +++ b/src/ros2_medkit_gateway/test/test_fault_manager.cpp @@ -30,9 +30,11 @@ #include "ros2_medkit_gateway/ros2/transports/ros2_fault_service_transport.hpp" #include "ros2_medkit_gateway/trigger_fault_subscriber.hpp" #include "ros2_medkit_msgs/msg/fault_event.hpp" +#include "ros2_medkit_msgs/srv/clear_fault.hpp" #include "ros2_medkit_msgs/srv/get_fault.hpp" #include "ros2_medkit_msgs/srv/get_rosbag.hpp" #include "ros2_medkit_msgs/srv/get_snapshots.hpp" +#include "ros2_medkit_msgs/srv/list_faults.hpp" using namespace std::chrono_literals; using ros2_medkit_gateway::FaultFailure; @@ -40,9 +42,11 @@ using ros2_medkit_gateway::FaultManager; using ros2_medkit_gateway::ResourceChange; using ros2_medkit_gateway::ResourceChangeNotifier; using ros2_medkit_gateway::TriggerFaultSubscriber; +using ros2_medkit_msgs::srv::ClearFault; using ros2_medkit_msgs::srv::GetFault; using ros2_medkit_msgs::srv::GetRosbag; using ros2_medkit_msgs::srv::GetSnapshots; +using ros2_medkit_msgs::srv::ListFaults; class FaultManagerTest : public ::testing::Test { protected: @@ -132,7 +136,7 @@ TEST_F(FaultManagerTest, GetSnapshotsServiceNotAvailable) { FaultManager fault_manager(std::make_shared(node_.get())); // Don't create a service, so it will timeout - auto result = fault_manager.get_snapshots("TEST_FAULT"); + auto result = fault_manager.get_snapshots("TEST_FAULT", "/cell_a/sensor"); EXPECT_FALSE(result.success); // On Humble, wait_for_service may report ready before DDS confirms absence, @@ -160,7 +164,7 @@ TEST_F(FaultManagerTest, GetSnapshotsSuccessWithValidJson) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_snapshots("MOTOR_OVERHEAT"); + auto result = fault_manager.get_snapshots("MOTOR_OVERHEAT", "/cell_a/sensor"); stop_spinning(); EXPECT_TRUE(result.success); @@ -195,7 +199,7 @@ TEST_F(FaultManagerTest, GetSnapshotsDrivesPrivateClientWithoutSpinningHostNode) // node_ is deliberately never spun by this test. FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_snapshots("SELF_DRIVEN_FAULT"); + auto result = fault_manager.get_snapshots("SELF_DRIVEN_FAULT", "/cell_a/sensor"); service_executor.cancel(); service_thread.join(); @@ -228,7 +232,7 @@ TEST_F(FaultManagerTest, GetSnapshotsUsesConfiguredFaultManagerNamespace) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_snapshots("NAMESPACED_FAULT"); + auto result = fault_manager.get_snapshots("NAMESPACED_FAULT", "/cell_a/sensor"); stop_spinning(); EXPECT_TRUE(result.success); @@ -258,7 +262,7 @@ TEST_F(FaultManagerTest, InvalidFaultManagerNamespaceFallsBackToRootServicePath) start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_snapshots("INVALID_NAMESPACE_FAULT"); + auto result = fault_manager.get_snapshots("INVALID_NAMESPACE_FAULT", "/cell_a/sensor"); stop_spinning(); EXPECT_TRUE(result.success); @@ -280,7 +284,7 @@ TEST_F(FaultManagerTest, GetSnapshotsSuccessWithTopicFilter) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - fault_manager.get_snapshots("TEST_FAULT", "/specific_topic"); + fault_manager.get_snapshots("TEST_FAULT", "/cell_a/sensor", "/specific_topic"); stop_spinning(); EXPECT_EQ(received_topic, "/specific_topic"); @@ -298,7 +302,7 @@ TEST_F(FaultManagerTest, GetSnapshotsErrorResponse) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_snapshots("NONEXISTENT_FAULT"); + auto result = fault_manager.get_snapshots("NONEXISTENT_FAULT", "/cell_a/sensor"); stop_spinning(); EXPECT_FALSE(result.success); @@ -317,7 +321,7 @@ TEST_F(FaultManagerTest, GetSnapshotsInvalidJsonFallback) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_snapshots("TEST_FAULT"); + auto result = fault_manager.get_snapshots("TEST_FAULT", "/cell_a/sensor"); stop_spinning(); EXPECT_TRUE(result.success); @@ -338,7 +342,7 @@ TEST_F(FaultManagerTest, GetSnapshotsEmptyResponse) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_snapshots("TEST_FAULT"); + auto result = fault_manager.get_snapshots("TEST_FAULT", "/cell_a/sensor"); stop_spinning(); EXPECT_TRUE(result.success); @@ -353,7 +357,7 @@ TEST_F(FaultManagerTest, GetRosbagServiceNotAvailable) { FaultManager fault_manager(std::make_shared(node_.get())); // Don't create a service, so it will timeout - auto result = fault_manager.get_rosbag("TEST_FAULT"); + auto result = fault_manager.get_rosbag("TEST_FAULT", "/cell_a/sensor"); EXPECT_FALSE(result.success); // On Humble, wait_for_service may report ready before DDS confirms absence, @@ -377,7 +381,7 @@ TEST_F(FaultManagerTest, GetRosbagSuccess) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_rosbag("TEST_ROSBAG_FAULT"); + auto result = fault_manager.get_rosbag("TEST_ROSBAG_FAULT", "/cell_a/sensor"); stop_spinning(); EXPECT_TRUE(result.success); @@ -411,7 +415,7 @@ TEST_F(FaultManagerTest, GetRosbagUsesConfiguredFaultManagerNamespace) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_rosbag("NAMESPACED_ROSBAG"); + auto result = fault_manager.get_rosbag("NAMESPACED_ROSBAG", "/cell_a/sensor"); stop_spinning(); EXPECT_TRUE(result.success); @@ -625,7 +629,7 @@ TEST_F(FaultManagerTest, GetRosbagNotFound) { start_spinning(); FaultManager fault_manager(std::make_shared(node_.get())); - auto result = fault_manager.get_rosbag("NONEXISTENT_FAULT"); + auto result = fault_manager.get_rosbag("NONEXISTENT_FAULT", "/cell_a/sensor"); stop_spinning(); EXPECT_FALSE(result.success); @@ -680,6 +684,141 @@ TEST_F(FaultManagerTest, GetFaultRefusedByTheStoreReportsDeclined) { EXPECT_EQ(result.error_message, "fault_code contains invalid character '~'"); } +// ============================================================================ +// The owning source reaches the wire +// +// A fault record is the pair (fault_code, owning reporting source), so every +// service that acts on one record carries the owner in its request. The +// gateway used to filter the answer instead, which meant asking the fault +// manager for an arbitrary record of that code and hoping it was the right +// one. These tests read the field off the request the service actually +// received, which is the only place that claim can be checked. +// ============================================================================ + +// @verifies REQ_INTEROP_013 +TEST_F(FaultManagerTest, GetFaultSendsTheOwningSourceInTheRequest) { + std::string received_source; + auto service = node_->create_service( + "/fault_manager/get_fault", [&received_source](const std::shared_ptr & request, + const std::shared_ptr & response) { + received_source = request->source_id; + response->success = true; + response->fault.fault_code = request->fault_code; + response->fault.reporting_sources = {request->source_id}; + }); + + start_spinning(); + FaultManager fault_manager(std::make_shared(node_.get())); + + auto result = fault_manager.get_fault_with_env("SHARED_CODE", "app_b"); + stop_spinning(); + + EXPECT_EQ(received_source, "app_b"); + ASSERT_TRUE(result.success) << result.error_message; + EXPECT_EQ(result.data["fault"]["source_id"], "app_b"); +} + +// @verifies REQ_INTEROP_015 +TEST_F(FaultManagerTest, ClearFaultSendsTheOwningSourceInTheRequest) { + std::string received_source; + std::string received_code; + auto service = node_->create_service( + "/fault_manager/clear_fault", + [&received_source, &received_code](const std::shared_ptr & request, + const std::shared_ptr & response) { + received_code = request->fault_code; + received_source = request->source_id; + response->success = true; + response->message = "cleared"; + }); + + start_spinning(); + FaultManager fault_manager(std::make_shared(node_.get())); + + auto result = fault_manager.clear_fault("SHARED_CODE", "app_a", /*skip_correlation_auto_clear=*/true); + stop_spinning(); + + EXPECT_TRUE(result.success) << result.error_message; + EXPECT_EQ(received_code, "SHARED_CODE"); + EXPECT_EQ(received_source, "app_a"); +} + +// @verifies REQ_INTEROP_088 +TEST_F(FaultManagerTest, GetSnapshotsSendsTheOwningSourceInTheRequest) { + std::string received_source; + auto service = node_->create_service( + "/fault_manager/get_snapshots", [&received_source](const std::shared_ptr & request, + const std::shared_ptr & response) { + received_source = request->source_id; + response->success = true; + response->data = "{}"; + }); + + start_spinning(); + FaultManager fault_manager(std::make_shared(node_.get())); + + fault_manager.get_snapshots("SHARED_CODE", "app_b", "/joint_states"); + stop_spinning(); + + EXPECT_EQ(received_source, "app_b"); +} + +// @verifies REQ_INTEROP_088 +TEST_F(FaultManagerTest, GetRosbagSendsTheOwningSourceInTheRequest) { + std::string received_source; + auto service = node_->create_service( + "/fault_manager/get_rosbag", [&received_source](const std::shared_ptr & request, + const std::shared_ptr & response) { + received_source = request->source_id; + response->success = true; + response->file_path = "/tmp/bag"; + }); + + start_spinning(); + FaultManager fault_manager(std::make_shared(node_.get())); + + fault_manager.get_rosbag("SHARED_CODE", "app_b"); + stop_spinning(); + + EXPECT_EQ(received_source, "app_b"); +} + +// A muted entry describes one record, so it has to name the owner. Without it +// the correlation view hands a client a bare fault_code, which addresses as many +// records as there are sources reporting it - the ambiguity the per-record +// identity exists to remove, reintroduced one layer up. +TEST_F(FaultManagerTest, MutedFaultEntriesNameTheRecordTheyDescribe) { + auto service = node_->create_service( + "/fault_manager/list_faults", + [](const std::shared_ptr & /*request*/, const std::shared_ptr & res) { + ros2_medkit_msgs::msg::MutedFaultInfo muted; + muted.fault_code = "SHARED_SYMPTOM"; + muted.root_cause_code = "ROOT"; + muted.rule_id = "cascade"; + muted.delay_ms = 150; + muted.source_id = "app_b"; + res->muted_faults.push_back(muted); + res->muted_count = 1; + }); + + start_spinning(); + FaultManager fault_manager(std::make_shared(node_.get())); + + auto result = fault_manager.list_faults("", /*include_prefailed=*/true, /*include_confirmed=*/true, + /*include_cleared=*/false, /*include_healed=*/false, + /*include_muted=*/true, /*include_clusters=*/false); + stop_spinning(); + + ASSERT_TRUE(result.success) << result.error_message; + ASSERT_TRUE(result.data.contains("muted_faults")); + ASSERT_EQ(result.data["muted_faults"].size(), 1u); + const auto & entry = result.data["muted_faults"][0]; + EXPECT_EQ(entry["fault_code"], "SHARED_SYMPTOM"); + EXPECT_EQ(entry["root_cause_code"], "ROOT"); + ASSERT_TRUE(entry.contains("source_id")) << "a muted entry must name the record's owner"; + EXPECT_EQ(entry["source_id"], "app_b"); +} + int main(int argc, char ** argv) { ::testing::InitGoogleTest(&argc, argv); return RUN_ALL_TESTS(); diff --git a/src/ros2_medkit_gateway/test/test_fault_manager_routing.cpp b/src/ros2_medkit_gateway/test/test_fault_manager_routing.cpp index ff001a933..7c0d4dc56 100644 --- a/src/ros2_medkit_gateway/test/test_fault_manager_routing.cpp +++ b/src/ros2_medkit_gateway/test/test_fault_manager_routing.cpp @@ -88,8 +88,10 @@ class MockFaultServiceTransport : public FaultServiceTransport { return r; } - FaultResult clear_fault(const std::string & fault_code, bool skip_correlation_auto_clear) override { + FaultResult clear_fault(const std::string & fault_code, const std::string & source_id, + bool skip_correlation_auto_clear) override { last_clear_code_ = fault_code; + last_clear_source_ = source_id; last_clear_skip_correlation_ = skip_correlation_auto_clear; ++clear_calls_; FaultResult r; @@ -99,8 +101,10 @@ class MockFaultServiceTransport : public FaultServiceTransport { return r; } - FaultResult get_snapshots(const std::string & fault_code, const std::string & topic) override { + FaultResult get_snapshots(const std::string & fault_code, const std::string & source_id, + const std::string & topic) override { last_snapshots_code_ = fault_code; + last_snapshots_source_ = source_id; last_snapshots_topic_ = topic; ++snapshots_calls_; FaultResult r; @@ -110,8 +114,9 @@ class MockFaultServiceTransport : public FaultServiceTransport { return r; } - FaultResult get_rosbag(const std::string & fault_code) override { - last_rosbag_code_ = fault_code; + FaultResult get_rosbag(const std::string & id, const std::string & source_id) override { + last_rosbag_code_ = id; + last_rosbag_source_ = source_id; ++rosbag_calls_; FaultResult r; r.success = rosbag_success_; @@ -181,6 +186,7 @@ class MockFaultServiceTransport : public FaultServiceTransport { json clear_data_ = json{{"success", true}, {"message", "ok"}}; std::string clear_error_; std::string last_clear_code_; + std::string last_clear_source_; bool last_clear_skip_correlation_ = false; int clear_calls_ = 0; @@ -188,6 +194,7 @@ class MockFaultServiceTransport : public FaultServiceTransport { json snapshots_data_ = json::object(); std::string snapshots_error_; std::string last_snapshots_code_; + std::string last_snapshots_source_; std::string last_snapshots_topic_; int snapshots_calls_ = 0; @@ -195,6 +202,7 @@ class MockFaultServiceTransport : public FaultServiceTransport { json rosbag_data_ = json::object(); std::string rosbag_error_; std::string last_rosbag_code_; + std::string last_rosbag_source_; int rosbag_calls_ = 0; bool list_rosbags_success_ = true; @@ -343,7 +351,7 @@ TEST(FaultManagerRoutingTest, ClearFaultDelegatesAndReturnsAutoClearedCodes) { mock->clear_data_ = {{"success", true}, {"message", "ok"}, {"auto_cleared_codes", json::array({"S1", "S2"})}}; FaultManager mgr(mock); - auto r = mgr.clear_fault("ROOT"); + auto r = mgr.clear_fault("ROOT", "/cell_a/root_node"); EXPECT_TRUE(r.success); EXPECT_EQ(mock->last_clear_code_, "ROOT"); @@ -351,6 +359,19 @@ TEST(FaultManagerRoutingTest, ClearFaultDelegatesAndReturnsAutoClearedCodes) { EXPECT_EQ(r.data["auto_cleared_codes"].size(), 2u); } +TEST(FaultManagerRoutingTest, ClearFaultCarriesTheOwningSourceToTheTransport) { + // A clear addresses one record, so the owner travels with the code. Without + // it the fault manager would have to guess which owner's record to clear + // whenever two sources report the same code. + auto mock = std::make_shared(); + FaultManager mgr(mock); + + mgr.clear_fault("SHARED_CODE", "app_b", /*skip_correlation_auto_clear=*/true); + + EXPECT_EQ(mock->last_clear_code_, "SHARED_CODE"); + EXPECT_EQ(mock->last_clear_source_, "app_b"); +} + TEST(FaultManagerRoutingTest, ClearFaultDefaultsToCorrelationAutoClear) { // Global `DELETE /api/v1/faults` and any unscoped caller should preserve // the legacy auto-clear-with-root behaviour (skip flag defaults to false). @@ -358,21 +379,21 @@ TEST(FaultManagerRoutingTest, ClearFaultDefaultsToCorrelationAutoClear) { mock->clear_success_ = true; FaultManager mgr(mock); - mgr.clear_fault("ROOT"); + mgr.clear_fault("ROOT", "/cell_a/root_node"); EXPECT_EQ(mock->last_clear_code_, "ROOT"); EXPECT_FALSE(mock->last_clear_skip_correlation_); } TEST(FaultManagerRoutingTest, ClearFaultForwardsSkipCorrelationFlag) { - // Per-entity DELETE routes call `clear_fault(code, /*skip=*/true)` so the - // fault manager does NOT cascade-clear correlated symptoms reported by + // Per-entity DELETE routes call `clear_fault(code, owner, /*skip=*/true)` so + // the fault manager does NOT cascade-clear correlated symptoms reported by // apps outside the addressed entity. Pin the routing here. auto mock = std::make_shared(); mock->clear_success_ = true; FaultManager mgr(mock); - mgr.clear_fault("ROOT", /*skip_correlation_auto_clear=*/true); + mgr.clear_fault("ROOT", "/cell_a/root_node", /*skip_correlation_auto_clear=*/true); EXPECT_EQ(mock->last_clear_code_, "ROOT"); EXPECT_TRUE(mock->last_clear_skip_correlation_); @@ -383,13 +404,26 @@ TEST(FaultManagerRoutingTest, GetSnapshotsRoutesTopicFilter) { mock->snapshots_data_ = {{"topics", json::object()}}; FaultManager mgr(mock); - auto r = mgr.get_snapshots("F1", "/joint_states"); + auto r = mgr.get_snapshots("F1", "/cell_a/sensor", "/joint_states"); EXPECT_TRUE(r.success); EXPECT_EQ(mock->last_snapshots_code_, "F1"); + EXPECT_EQ(mock->last_snapshots_source_, "/cell_a/sensor"); EXPECT_EQ(mock->last_snapshots_topic_, "/joint_states"); } +TEST(FaultManagerRoutingTest, GetRosbagCarriesTheOwningSourceToTheTransport) { + // Snapshots and recordings belong to one record. The owner scopes the + // fault-code lookup so a shared code does not serve another owner's bytes. + auto mock = std::make_shared(); + FaultManager mgr(mock); + + mgr.get_rosbag("SHARED_CODE", "app_b"); + + EXPECT_EQ(mock->last_rosbag_code_, "SHARED_CODE"); + EXPECT_EQ(mock->last_rosbag_source_, "app_b"); +} + TEST(FaultManagerRoutingTest, GetRosbagPropagatesErrorMessage) { auto mock = std::make_shared(); mock->rosbag_success_ = false; diff --git a/src/ros2_medkit_gateway/test/test_fault_trigger_engine.cpp b/src/ros2_medkit_gateway/test/test_fault_trigger_engine.cpp index e54415a9d..905a4912d 100644 --- a/src/ros2_medkit_gateway/test/test_fault_trigger_engine.cpp +++ b/src/ros2_medkit_gateway/test/test_fault_trigger_engine.cpp @@ -145,6 +145,25 @@ TEST(FaultTriggerEngineTest, CreateWithoutEnumeratorSkipsExistenceCheck) { EXPECT_TRUE(engine.create("tank", make_body("whatever", ">", 1.0, "F", "ERROR"))); } +// A code another app's rule holds is refused, and the refusal states the +// engine's own rule. The fault manager would keep the two apps' records apart, +// so a message blaming the fault store sends the operator after the wrong cause. +TEST(FaultTriggerEngineTest, ADuplicateCodeIsRefusedWithTheEnginesOwnRule) { + FaultTriggerEngine engine("", nullptr, nullptr, nullptr, nullptr); + auto first = engine.create("app_a", make_body("x", ">", 1.0, "SHARED", "ERROR")); + ASSERT_TRUE(first); + + auto second = engine.create("app_b", make_body("y", ">", 1.0, "SHARED", "ERROR")); + + ASSERT_FALSE(second); + EXPECT_EQ(second.error().first, 409); + const auto & msg = second.error().second; + EXPECT_NE(msg.find("rule '" + first->id + "' on app 'app_a'"), std::string::npos) << msg; + EXPECT_NE(msg.find("The trigger engine keeps one rule per fault code across every app"), std::string::npos) << msg; + EXPECT_EQ(msg.find("fault store"), std::string::npos) << "the refusal is the engine's rule, not the store's: " << msg; + EXPECT_TRUE(engine.list("app_b").empty()); +} + TEST(FaultTriggerEngineTest, ListIsScopedPerApp) { FaultTriggerEngine engine("", nullptr, nullptr, nullptr, nullptr); ASSERT_TRUE(engine.create("app_a", make_body("x", ">", 1.0, "FA", "ERROR"))); diff --git a/src/ros2_medkit_gateway/test/test_handler_context.cpp b/src/ros2_medkit_gateway/test/test_handler_context.cpp index 968f29469..05f99f037 100644 --- a/src/ros2_medkit_gateway/test/test_handler_context.cpp +++ b/src/ros2_medkit_gateway/test/test_handler_context.cpp @@ -40,6 +40,7 @@ #include "ros2_medkit_gateway/core/models/error_info.hpp" #include "ros2_medkit_gateway/core/models/thread_safe_entity_cache.hpp" #include "ros2_medkit_gateway/gateway_node.hpp" +#include "ros2_medkit_gateway/http/handlers/fault_handlers.hpp" #include "ros2_medkit_gateway/http/handlers/handler_context.hpp" #include "ros2_medkit_gateway/http/typed_router.hpp" @@ -1797,6 +1798,76 @@ TEST(ResolveEntitySourceFqnsTest, ExternalAppWithStrayRosBindingOwnsFaultsByBare EXPECT_EQ(fqns, std::set{"process"}); } +// ============================================================================= +// build_source_entity_map - whose lock the global fault clear consults +// +// `DELETE /api/v1/faults` skips a record whose owning entity is locked by +// another client. It finds that entity by looking the record's reporting source +// up in this map, so a source the map cannot attribute is a lock the route +// silently walks through. +// ============================================================================= +TEST(BuildSourceEntityMapTest, BoundAppIsFoundByItsFqn) { + ThreadSafeEntityCache cache; + cache.update_apps({make_app_with_binding("temp_sensor", "temp_sensor", "/powertrain/engine")}); + + const auto map = handlers::FaultHandlers::build_source_entity_map(cache); + + ASSERT_EQ(map.count("/powertrain/engine/temp_sensor"), 1u); + EXPECT_EQ(map.at("/powertrain/engine/temp_sensor"), "temp_sensor"); +} + +// An external app has no ROS binding, so its effective_fqn() is empty and the +// old map held nothing for it. It reports under its bare id, which is the value +// that arrives on the wire, so the lock of every plugin-owned app went +// unhonoured on this route. +TEST(BuildSourceEntityMapTest, ExternalAppIsFoundByItsBareId) { + ThreadSafeEntityCache cache; + cache.update_apps({make_external_app("plc_line1", "plc_hw")}); + + const auto map = handlers::FaultHandlers::build_source_entity_map(cache); + + ASSERT_EQ(map.count("plc_line1"), 1u) << "an external app's bare id must resolve to its entity"; + EXPECT_EQ(map.at("plc_line1"), "plc_line1"); +} + +// A protocol bridge raises its link faults under the component's own id, and +// the map was built from apps only, so it never held a component at all. +TEST(BuildSourceEntityMapTest, ExternalComponentIsFoundByItsBareId) { + ThreadSafeEntityCache cache; + Component device; + device.id = "line_controller"; + device.external = true; + cache.update_components({device}); + + const auto map = handlers::FaultHandlers::build_source_entity_map(cache); + + ASSERT_EQ(map.count("line_controller"), 1u) << "an external component's bare id must resolve to its entity"; + EXPECT_EQ(map.at("line_controller"), "line_controller"); +} + +// The #516 guardrail, restated for the lock map: a runtime host component never +// claims its bare id, so attributing a record to it would let an unrelated +// lock block a clear. +TEST(BuildSourceEntityMapTest, InternalComponentClaimsNoBareId) { + ThreadSafeEntityCache cache; + Component host; + host.id = "runtime_host"; + cache.update_components({host}); + + EXPECT_EQ(handlers::FaultHandlers::build_source_entity_map(cache).count("runtime_host"), 0u); +} + +// An unbound non-external app owns no reporting source at all: granting it its +// bare id would let it claim records it never reported. +TEST(BuildSourceEntityMapTest, UnboundInternalAppContributesNothing) { + ThreadSafeEntityCache cache; + App unbound; + unbound.id = "manifest_only"; + cache.update_apps({unbound}); + + EXPECT_TRUE(handlers::FaultHandlers::build_source_entity_map(cache).empty()); +} + TEST(ResolveEntitySourceFqnsTest, ExternalComponentWithNoAppsOwnsFaultsUnderItsOwnId) { // The #516 case: a PLC modeled as an external Component with no child App. // It reports faults under its own component id; the scope must be exactly diff --git a/src/ros2_medkit_gateway/test/test_plugin_context_aggregation.cpp b/src/ros2_medkit_gateway/test/test_plugin_context_aggregation.cpp index 9119cd2f1..f9622fd5e 100644 --- a/src/ros2_medkit_gateway/test/test_plugin_context_aggregation.cpp +++ b/src/ros2_medkit_gateway/test/test_plugin_context_aggregation.cpp @@ -179,15 +179,15 @@ class FakeFaultTransport : public ros2_medkit_gateway::FaultServiceTransport { const std::string & /*source_id*/) override { return {false, json::object(), "not implemented"}; } - ros2_medkit_gateway::FaultResult clear_fault(const std::string & /*fault_code*/, + ros2_medkit_gateway::FaultResult clear_fault(const std::string & /*fault_code*/, const std::string & /*source_id*/, bool /*skip_correlation_auto_clear*/) override { return {false, json::object(), "not implemented"}; } - ros2_medkit_gateway::FaultResult get_snapshots(const std::string & /*fault_code*/, + ros2_medkit_gateway::FaultResult get_snapshots(const std::string & /*fault_code*/, const std::string & /*source_id*/, const std::string & /*topic*/) override { return {false, json::object(), "not implemented"}; } - ros2_medkit_gateway::FaultResult get_rosbag(const std::string & /*fault_code*/) override { + ros2_medkit_gateway::FaultResult get_rosbag(const std::string & /*id*/, const std::string & /*source_id*/) override { return {false, json::object(), "not implemented"}; } ros2_medkit_gateway::FaultResult list_rosbags(const std::string & /*entity_fqn*/) override { diff --git a/src/ros2_medkit_gateway/test/test_plugin_entity_routing.cpp b/src/ros2_medkit_gateway/test/test_plugin_entity_routing.cpp index 260fad967..20e8b286a 100644 --- a/src/ros2_medkit_gateway/test/test_plugin_entity_routing.cpp +++ b/src/ros2_medkit_gateway/test/test_plugin_entity_routing.cpp @@ -14,6 +14,10 @@ #include +#include +#include + +#include "ros2_medkit_gateway/core/plugins/plugin_http_types.hpp" #include "ros2_medkit_gateway/core/plugins/plugin_manager.hpp" #include "ros2_medkit_gateway/core/providers/data_provider.hpp" #include "ros2_medkit_gateway/core/providers/fault_provider.hpp" @@ -21,6 +25,7 @@ #include "ros2_medkit_gateway/dto/data.hpp" #include "ros2_medkit_gateway/dto/faults.hpp" #include "ros2_medkit_gateway/dto/operations.hpp" +#include "ros2_medkit_gateway/entity_freeze_frame_capture.hpp" using namespace ros2_medkit_gateway; // json alias already available via ros2_medkit_gateway namespace headers @@ -129,6 +134,237 @@ class MockMixedFitnessPlugin : public GatewayPlugin, public DataProvider, public } }; +// A plugin that serves /data ONLY through its DataProvider: no vendor route, +// which is what an in-tree plugin looks like. +class ProviderOnlyPlugin : public GatewayPlugin, public DataProvider { + public: + std::string name() const override { + return "provider_only"; + } + void configure(const json & /*config*/) override { + } + void shutdown() override { + } + + tl::expected list_data(const std::string & entity_id) override { + if (throws_) { + throw std::runtime_error("provider exploded"); + } + return dto::DataListResult{ + json{{"connected", true}, {"items", json::array({{{"id", "level"}, {"value", 7.5}, {"entity", entity_id}}})}}}; + } + tl::expected read_data(const std::string & /*entity_id*/, + const std::string & resource) override { + return dto::DataValue{json{{"value", resource}}}; + } + tl::expected + write_data(const std::string & /*entity_id*/, const std::string & /*resource*/, const json & /*payload*/) override { + return dto::DataWriteResult{json{{"status", "ok"}}}; + } + + bool throws_ = false; +}; + +// A plugin that serves the SAME entity through BOTH a DataProvider and its own +// data route, with different values in each. Without this fixture the +// provider-first order is unpinned: every provider-only and route-only case +// passes under either order, so restoring route-first breaks nothing and is +// caught by nothing. +class ProviderAndRoutePlugin : public GatewayPlugin, public DataProvider { + public: + static constexpr double kProviderLevel = 11.0; + static constexpr double kRouteLevel = 99.0; + + std::string name() const override { + return "provider_and_route"; + } + void configure(const json & /*config*/) override { + } + void shutdown() override { + } + + std::vector get_routes() override { + return { + {"GET", R"(apps/([^/]+)/x-plc-data)", + [this](const PluginRequest & /*req*/, PluginResponse & res) { + ++route_calls; + res.send_json(json{{"connected", route_connected}, + {"items", json::array({{{"id", "level"}, {"value", kRouteLevel}}})}}); + }}, + }; + } + + tl::expected list_data(const std::string & /*entity_id*/) override { + ++provider_calls; + return dto::DataListResult{json{{"connected", provider_connected}, + {"items", json::array({{{"id", "level"}, {"value", kProviderLevel}}})}}}; + } + tl::expected read_data(const std::string & /*entity_id*/, + const std::string & resource) override { + return dto::DataValue{json{{"value", resource}}}; + } + tl::expected + write_data(const std::string & /*entity_id*/, const std::string & /*resource*/, const json & /*payload*/) override { + return dto::DataWriteResult{json{{"status", "ok"}}}; + } + + int provider_calls = 0; + int route_calls = 0; + bool provider_connected = true; + bool route_connected = true; +}; + +// The same both-sources shape, with a provider that throws. Used to show the +// route really is reachable on this fixture, so the order test above is about +// the order and not about an unreachable route. +class ThrowingProviderAndRoutePlugin : public ProviderAndRoutePlugin { + public: + std::string name() const override { + return "throwing_provider_and_route"; + } + tl::expected list_data(const std::string & /*entity_id*/) override { + ++provider_calls; + throw std::runtime_error("provider exploded"); + } +}; + +// A plugin that exposes no DataProvider at all and registers no route either, +// which is every grouping-only plugin. +class NoDataPlugin : public GatewayPlugin { + public: + std::string name() const override { + return "no_data"; + } + void configure(const json & /*config*/) override { + } + void shutdown() override { + } +}; + +// ============================================================================= +// fetch_entity_data_content - which source an entity's /data comes from +// +// Provider first, the plugin's own vendor route only when it exposes no +// provider. The fault-trigger engine reads rule values through this and +// enumerates a rule's data points through it, so reading the route first meant +// a rule on a provider-served app evaluated to nothing every tick and sat +// there silently. +// ============================================================================= + +TEST(PluginEntityDataContent, ProviderServesAnEntityWithNoVendorRoute) { + PluginManager mgr; + auto plugin = std::make_unique(); + mgr.add_plugin(std::move(plugin)); + mgr.register_entity_ownership("provider_only", {"tank"}); + + auto content = mgr.fetch_entity_data_content("tank"); + + ASSERT_TRUE(content.has_value()) << "an entity served by a DataProvider must not read as having no data"; + ASSERT_TRUE(content->contains("items")); + EXPECT_EQ((*content)["items"][0]["id"], "level"); + EXPECT_DOUBLE_EQ((*content)["items"][0]["value"].get(), 7.5); +} + +TEST(PluginEntityDataContent, AThrowingProviderFallsThroughInsteadOfEscaping) { + // The fault-trigger engine calls this on its own evaluation loop, and a + // plugin exception must not leave that loop. + PluginManager mgr; + auto plugin = std::make_unique(); + plugin->throws_ = true; + mgr.add_plugin(std::move(plugin)); + mgr.register_entity_ownership("provider_only", {"tank"}); + + std::optional content; + EXPECT_NO_THROW(content = mgr.fetch_entity_data_content("tank")); + // No vendor route behind it, so there is nothing left to answer with. + EXPECT_FALSE(content.has_value()); +} + +TEST(PluginEntityDataContent, NeitherProviderNorRouteIsNullopt) { + PluginManager mgr; + mgr.add_plugin(std::make_unique()); + mgr.register_entity_ownership("no_data", {"grouping"}); + + EXPECT_FALSE(mgr.fetch_entity_data_content("grouping").has_value()); +} + +TEST(PluginEntityDataContent, AnUnownedEntityIsNullopt) { + PluginManager mgr; + EXPECT_FALSE(mgr.fetch_entity_data_content("nobody").has_value()); +} + +// One entity, both sources, different values in each. This is the only fixture +// in which the ORDER is observable at all. +TEST(PluginEntityDataContent, TheProviderIsReadBeforeTheVendorRoute) { + PluginManager mgr; + auto plugin = std::make_unique(); + auto * raw = plugin.get(); + mgr.add_plugin(std::move(plugin)); + mgr.register_entity_ownership("provider_and_route", {"tank"}); + + auto content = mgr.fetch_entity_data_content("tank"); + + ASSERT_TRUE(content.has_value()); + EXPECT_DOUBLE_EQ((*content)["items"][0]["value"].get(), ProviderAndRoutePlugin::kProviderLevel) + << "the route answered first, so a provider-served entity reads its vendor route instead"; + EXPECT_EQ(raw->provider_calls, 1); + EXPECT_EQ(raw->route_calls, 0) << "the route must not be dispatched when the provider answered"; +} + +// The route is the fallback, not the second opinion. On the same both-sources +// fixture, a provider that throws hands over and the value then comes from the +// route, which is what makes the previous test about ORDER rather than about the +// route being unreachable. +TEST(PluginEntityDataContent, TheVendorRouteAnswersWhenTheProviderThrows) { + PluginManager mgr; + auto plugin = std::make_unique(); + auto * raw = plugin.get(); + mgr.add_plugin(std::move(plugin)); + mgr.register_entity_ownership("throwing_provider_and_route", {"tank"}); + + auto content = mgr.fetch_entity_data_content("tank"); + + ASSERT_TRUE(content.has_value()); + EXPECT_DOUBLE_EQ((*content)["items"][0]["value"].get(), ProviderAndRoutePlugin::kRouteLevel); + EXPECT_EQ(raw->route_calls, 1) << "the route is the fallback and must have been dispatched"; +} + +// The trigger fetcher's link-down guard reads `connected` off this content. A +// provider envelope without it left the guard unable to fire for every +// provider-served entity, so a rule evaluated on frozen last-known values for a +// whole outage. +TEST(PluginEntityDataContent, ADisconnectedProviderEnvelopeReportsDisconnected) { + PluginManager mgr; + auto plugin = std::make_unique(); + auto * raw = plugin.get(); + mgr.add_plugin(std::move(plugin)); + mgr.register_entity_ownership("provider_and_route", {"tank"}); + raw->provider_connected = false; + + auto content = mgr.fetch_entity_data_content("tank"); + + ASSERT_TRUE(content.has_value()); + EXPECT_TRUE(EntityFreezeFrameCapture::content_has_live_data(*content)); + EXPECT_TRUE(EntityFreezeFrameCapture::content_reports_disconnected(*content)) + << "a provider envelope reporting a down link must read as disconnected, " + "or the trigger fetcher evaluates on frozen values"; +} + +// The control on the same harness: a connected provider envelope must NOT read +// as disconnected, so the assertion above is about the flag and not about the +// guard answering true for everything. +TEST(PluginEntityDataContent, AConnectedProviderEnvelopeReportsConnected) { + PluginManager mgr; + auto plugin = std::make_unique(); + mgr.add_plugin(std::move(plugin)); + mgr.register_entity_ownership("provider_and_route", {"tank"}); + + auto content = mgr.fetch_entity_data_content("tank"); + + ASSERT_TRUE(content.has_value()); + EXPECT_FALSE(EntityFreezeFrameCapture::content_reports_disconnected(*content)); +} + // ============================================================================= // Entity Ownership Tests // ============================================================================= diff --git a/src/ros2_medkit_gateway/test/test_sse_fault_handler.cpp b/src/ros2_medkit_gateway/test/test_sse_fault_handler.cpp index c51d19fb2..6ed69e384 100644 --- a/src/ros2_medkit_gateway/test/test_sse_fault_handler.cpp +++ b/src/ros2_medkit_gateway/test/test_sse_fault_handler.cpp @@ -29,6 +29,7 @@ #include "ros2_medkit_gateway/core/config.hpp" #include "ros2_medkit_gateway/core/discovery/models/app.hpp" +#include "ros2_medkit_gateway/core/discovery/models/component.hpp" #include "ros2_medkit_gateway/core/http/sse_client_tracker.hpp" #include "ros2_medkit_gateway/core/models/thread_safe_entity_cache.hpp" #include "ros2_medkit_gateway/fault_manager_paths.hpp" @@ -496,6 +497,45 @@ TEST_F(SSEFaultHandlerTest, CoalescedReplayKeepsTransitionsAndEndsOnCurrentState release_stream(res); } +TEST_F(SSEFaultHandlerTest, AnotherOwnersEventsDoNotSupersedeARecordsLastUpdate) { + // A fault code names as many records as there are sources reporting it. A + // newer event of one owner's record says nothing about another owner's record + // of the same code, so it must not coalesce that record's last update away: + // a lagging client would never learn that record's current state, and the + // loss would not even be counted. + auto req = make_stream_request("127.0.0.1"); + httplib::Response res; + handler_->handle_stream(req, res); // cursor open, never drained while the buffer fills + + auto owned_by = [](FaultEvent event, const std::string & owner) { + event.fault.reporting_sources = {owner}; + return event; + }; + enqueue_event(owned_by(make_fault_event(FaultEvent::EVENT_CONFIRMED, "SHARED", 1), "/owner_a")); + enqueue_event(owned_by(make_fault_event(FaultEvent::EVENT_UPDATED, "SHARED", 2), "/owner_a")); + enqueue_event(owned_by(make_fault_event(FaultEvent::EVENT_CONFIRMED, "SHARED", 3), "/owner_b")); + for (int i = 1; i <= 148; ++i) { + enqueue_event(owned_by(make_fault_event(FaultEvent::EVENT_UPDATED, "SHARED", 100 + i), "/owner_b")); + } + + EXPECT_GT(handler_->coalesced_events(), 0u) << "owner_b's own superseded updates are the ones to coalesce"; + EXPECT_EQ(handler_->dropped_events(), 0u); + + auto output = read_stream_once(res, 100); + int owner_a_updates = 0; + for (auto pos = output.find("event: "); pos != std::string::npos; pos = output.find("event: ", pos + 1)) { + auto payload = parse_sse_payload(output.substr(pos)); + if (payload["event_type"] == "fault_updated" && + payload["fault"]["reporting_sources"] == json::array({"/owner_a"})) { + ++owner_a_updates; + EXPECT_DOUBLE_EQ(payload["timestamp"].get(), 2.0); + } + } + EXPECT_EQ(owner_a_updates, 1) << "owner_a's last update was coalesced away by owner_b's events of the same code"; + + release_stream(res); +} + TEST_F(SSEFaultHandlerTest, OwedDistinctEventsLostUnderPressureAreCountedAndLogged) { // The counter this PR is about: a live client is owed events, the buffer // overflows with distinct fault codes (nothing to coalesce), so events are @@ -816,6 +856,91 @@ TEST_F(SSEFaultHandlerTest, StreamResolvesRuntimeCollisionRenamedApp) { release_stream(res); } +// An external app owns its record under its bare SOVD id, which is not a ROS +// node FQN. An entity id carries no slash, so the last-segment fallback looks +// it up as an app id unchanged - this pins that, since the bare-id case is +// what the record model puts on the stream and nothing else covers it. +TEST_F(SSEFaultHandlerTest, StreamResolvesABareExternalAppId) { + App app; + app.id = "plc_process"; + app.name = "process"; + app.source = "plugin"; + app.external = true; + auto & cache = const_cast(node_->get_thread_safe_cache()); + cache.update_apps({app}); + + auto event = make_fault_event(FaultEvent::EVENT_CONFIRMED, "PROCESS_LEVEL_HIGH", 72); + event.fault.reporting_sources = {"plc_process"}; + enqueue_event(event); + + auto req = make_stream_request("127.0.0.1"); + httplib::Response res; + handler_->handle_stream(req, res); + + auto payload = parse_sse_payload(read_stream_once(res, 1)); + + ASSERT_TRUE(payload.contains("x-medkit")) << payload.dump(); + EXPECT_EQ(payload["x-medkit"]["entity_type"], "apps"); + EXPECT_EQ(payload["x-medkit"]["entity_id"], "plc_process"); + + release_stream(res); +} + +// A protocol bridge raises its link faults under the component's own id. The +// hint must name the component, because that is the entity whose fault routes +// serve the record. +TEST_F(SSEFaultHandlerTest, StreamResolvesABareExternalComponentId) { + ros2_medkit_gateway::Component component; + component.id = "line_controller"; + component.name = "Line Controller"; + component.source = "plugin"; + component.external = true; + auto & cache = const_cast(node_->get_thread_safe_cache()); + cache.update_all({}, {component}, {}, {}); + + auto event = make_fault_event(FaultEvent::EVENT_CONFIRMED, "DEVICE_COMMS_LOST", 73); + event.fault.reporting_sources = {"line_controller"}; + enqueue_event(event); + + auto req = make_stream_request("127.0.0.1"); + httplib::Response res; + handler_->handle_stream(req, res); + + auto payload = parse_sse_payload(read_stream_once(res, 1)); + + ASSERT_TRUE(payload.contains("x-medkit")) << payload.dump(); + EXPECT_EQ(payload["x-medkit"]["entity_type"], "components"); + EXPECT_EQ(payload["x-medkit"]["entity_id"], "line_controller"); + + release_stream(res); +} + +// Only an external component claims its bare id as a reporting source. A +// runtime host component never does, and naming it would point the consumer at +// an entity whose own fault routes drop the record. +TEST_F(SSEFaultHandlerTest, StreamDoesNotNameANonExternalComponent) { + ros2_medkit_gateway::Component component; + component.id = "runtime_host"; + component.name = "runtime host"; + component.source = "heuristic"; + auto & cache = const_cast(node_->get_thread_safe_cache()); + cache.update_all({}, {component}, {}, {}); + + auto event = make_fault_event(FaultEvent::EVENT_CONFIRMED, "HOST_FAULT", 74); + event.fault.reporting_sources = {"runtime_host"}; + enqueue_event(event); + + auto req = make_stream_request("127.0.0.1"); + httplib::Response res; + handler_->handle_stream(req, res); + + auto payload = parse_sse_payload(read_stream_once(res, 1)); + + EXPECT_FALSE(payload.contains("x-medkit")) << payload.dump(); + + release_stream(res); +} + TEST_F(SSEFaultHandlerTest, StreamOmitsXMedkitWhenReportingSourcesEmpty) { auto event = make_fault_event(FaultEvent::EVENT_CONFIRMED, "ORPHAN", 40); event.fault.reporting_sources.clear(); diff --git a/src/ros2_medkit_integration_tests/test/features/test_faults_owner_identity.test.py b/src/ros2_medkit_integration_tests/test/features/test_faults_owner_identity.test.py new file mode 100644 index 000000000..e43b25e49 --- /dev/null +++ b/src/ros2_medkit_integration_tests/test/features/test_faults_owner_identity.test.py @@ -0,0 +1,499 @@ +#!/usr/bin/env python3 +# Copyright 2026 bburda +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""Two owners of one fault code, over the whole HTTP stack. + +A fault record is the pair (``fault_code``, reporting source). The source is the +``source_id`` the reporter used and it owns the record, so two apps reporting one +code are two records with their own status, timestamps and clear. + +The unit layers cover the pieces separately: the record selection and its 404 / +409 mapping (``SelectScopedFaultTest``), the scope resolution +(``ResolveEntitySourceFqnsTest``), the wire fields (``FaultListItemSchema``) and +the owner reaching the service request (``FaultManagerTest.*SendsTheOwningSource*``). +None of them runs a real fault manager behind a real gateway, which is where the +model either holds or collapses back to one record per code. + +The topology is two external apps under one component, so both records land in +one component's fault scope. What this pins: + +* ``GET /faults`` lists two items with the same code and distinct ``source_id``. +* Each app's own detail route serves its own record and nothing else. +* ``DELETE`` on one app's record leaves the other app's record CONFIRMED. Under + a code-keyed store one clear took both. +* The component owns both, so its per-code detail route names neither: ``409`` + ``x-medkit-ambiguous-fault`` listing both owners, instead of serving whichever + record the store happened to list first. +* ``DELETE /components//faults`` clears both, one record at a time. +* Every stream frame carries the entity hint for the record it describes, and + the two records resolve to their own owners. +* The served OpenAPI declares what the routes now do. + +@verifies REQ_INTEROP_012 +@verifies REQ_INTEROP_013 +@verifies REQ_INTEROP_014 +@verifies REQ_INTEROP_015 +""" + +import json +import os +import threading +import time +import unittest + +from ament_index_python.packages import get_package_share_directory +import launch_testing +import rclpy +from rclpy.node import Node +import requests +from ros2_medkit_msgs.msg import Fault +from ros2_medkit_msgs.srv import ReportFault + +from ros2_medkit_test_utils.constants import ALLOWED_EXIT_CODES +from ros2_medkit_test_utils.gateway_test_case import GatewayTestCase +from ros2_medkit_test_utils.launch_helpers import create_test_launch + + +OWNER_A = 'owner-a' +OWNER_B = 'owner-b' +HOST_COMPONENT = 'shared-code-hub' +SHARED_CODE = 'SHARED_OVERFLOW' +PRIME_CODE = 'OWNER_IDENTITY_PRIME' +FAULT_TIMEOUT = 30.0 + + +def generate_test_description(): + manifest_path = os.path.join( + get_package_share_directory('ros2_medkit_gateway'), + 'config', 'examples', 'fault_owner_identity_manifest.yaml', + ) + return create_test_launch( + demo_nodes=[], + fault_manager=True, + # Negative threshold: a record confirms after a couple of FAILED events, + # and each owner debounces on its own reports. + fault_manager_params={'confirmation_threshold': -2}, + gateway_params={ + 'discovery.mode': 'hybrid', + 'discovery.manifest_path': manifest_path, + 'discovery.manifest_strict_validation': False, + # Locking off, so the per-code DELETE publishes no lock contract and + # the only 409 its document can carry is the ambiguous-fault one the + # route declares itself. With locking on, the lock marker adds a 409 + # of its own and the DELETE check in test_07 could never fail. + # Nothing else here takes a lock. + 'locking.enabled': False, + }, + ) + + +class TestFaultsOwnerIdentity(GatewayTestCase): + """Two sources reporting one code are two independently addressable records.""" + + MIN_EXPECTED_APPS = 2 + REQUIRED_APPS = {OWNER_A, OWNER_B} + + @classmethod + def setUpClass(cls): + rclpy.init() + cls._reporter = Node('fault_owner_identity_reporter') + cls._report_client = cls._reporter.create_client( + ReportFault, '/fault_manager/report_fault' + ) + super().setUpClass() + assert cls._report_client.wait_for_service(timeout_sec=15.0), \ + 'report_fault service not available' + + @classmethod + def tearDownClass(cls): + cls._reporter.destroy_node() + rclpy.shutdown() + + def _report(self, source_id, *, fault_code=SHARED_CODE, + event_type=ReportFault.Request.EVENT_FAILED, times=4): + """Report under one owner, fire and forget. + + Same shape as the production fault_reporter client: the reply path is + not discovery-matched by wait_for_service, so an early round trip can + lose its response even though the record was written. The behavioural + assertion is the HTTP poll, which retries. + """ + for _ in range(times): + req = ReportFault.Request() + req.fault_code = fault_code + req.event_type = event_type + req.severity = Fault.SEVERITY_ERROR + req.description = f'reported by {source_id}' + req.source_id = source_id + self._report_client.call_async(req) + rclpy.spin_once(self._reporter, timeout_sec=0.1) + + def _raise_both(self): + self._report(OWNER_A) + self._report(OWNER_B) + self.wait_for_fault(f'/apps/{OWNER_A}', SHARED_CODE, max_wait=FAULT_TIMEOUT) + self.wait_for_fault(f'/apps/{OWNER_B}', SHARED_CODE, max_wait=FAULT_TIMEOUT) + + def _records_of_shared_code(self): + body = self.get_json('/faults') + items = body.get('items', body.get('faults', [])) + return [item for item in items if item.get('fault_code') == SHARED_CODE] + + def _status_on(self, app_id): + """Status of that app's own record, or None when it holds none.""" + body = self.get_json(f'/apps/{app_id}/faults?status=all') + for item in body.get('items', []): + if item.get('fault_code') == SHARED_CODE: + return item.get('status') + return None + + # Method order is alphabetical, and the clearing cases consume the records + # the reading cases assert on, so the names carry the order. + + def test_01_the_global_list_shows_one_item_per_owner(self): + """One code, two sources, two items - each naming its own owner. + + @verifies REQ_INTEROP_012 + """ + self._raise_both() + + records = self._records_of_shared_code() + + self.assertEqual( + len(records), 2, + f'two sources reporting {SHARED_CODE} must be two records, got: {records}', + ) + owners = sorted(r.get('source_id') for r in records) + self.assertEqual(owners, [OWNER_A, OWNER_B]) + for record in records: + self.assertEqual( + record.get('reporting_sources'), [record.get('source_id')], + 'a record names exactly its own owner', + ) + + def test_02_each_app_detail_serves_its_own_record(self): + """The app route addresses the record that app owns. + + @verifies REQ_INTEROP_013 + """ + self._raise_both() + + for owner in (OWNER_A, OWNER_B): + detail = self.get_json(f'/apps/{owner}/faults/{SHARED_CODE}') + self.assertEqual(detail['item']['code'], SHARED_CODE) + self.assertEqual( + detail['x-medkit']['owner'], owner, + f"/apps/{owner} served another owner's record", + ) + self.assertNotIn( + 'source_id', detail['x-medkit'], + "the detail must not reuse the list's source_id key for the owner", + ) + self.assertEqual(detail['x-medkit']['reporting_sources'], [owner]) + + def test_03_the_component_detail_refuses_to_pick_an_owner(self): + """Both records are in the component's scope, so the code names neither. + + @verifies REQ_INTEROP_013 + """ + self._raise_both() + + response = requests.get( + f'{self.BASE_URL}/components/{HOST_COMPONENT}/faults/{SHARED_CODE}', + timeout=10, + ) + + self.assertEqual(response.status_code, 409, response.text) + body = response.json() + self.assertEqual(body.get('vendor_code', body.get('error_code')), + 'x-medkit-ambiguous-fault', body) + owners = sorted(body.get('parameters', body).get('owners', [])) + self.assertEqual(owners, [OWNER_A, OWNER_B], body) + + def test_04_the_component_list_still_shows_both(self): + """Ambiguous to address by code is not invisible: the list has both. + + @verifies REQ_INTEROP_012 + """ + self._raise_both() + + body = self.get_json(f'/components/{HOST_COMPONENT}/faults') + owners = sorted( + item.get('source_id') for item in body.get('items', []) + if item.get('fault_code') == SHARED_CODE + ) + + self.assertEqual(owners, [OWNER_A, OWNER_B], body) + + def test_05_clearing_one_owners_record_leaves_the_other(self): + """The clear addresses one record. Under a code key it took both. + + @verifies REQ_INTEROP_015 + """ + self._raise_both() + + self.delete_request(f'/apps/{OWNER_A}/faults/{SHARED_CODE}') + + deadline = time.monotonic() + FAULT_TIMEOUT + while time.monotonic() < deadline: + if self._status_on(OWNER_A) == 'CLEARED': + break + time.sleep(0.2) + self.assertEqual(self._status_on(OWNER_A), 'CLEARED') + self.assertEqual( + self._status_on(OWNER_B), 'CONFIRMED', + f"clearing {OWNER_A}'s record must not touch {OWNER_B}'s record of the same code", + ) + + def test_06_the_component_bulk_clear_takes_every_record(self): + """Per-entity DELETE clears each in-scope record with its own owner. + + @verifies REQ_INTEROP_014 + """ + self._raise_both() + + self.delete_request(f'/components/{HOST_COMPONENT}/faults') + + deadline = time.monotonic() + FAULT_TIMEOUT + while time.monotonic() < deadline: + if (self._status_on(OWNER_A) == 'CLEARED' + and self._status_on(OWNER_B) == 'CLEARED'): + break + time.sleep(0.2) + self.assertEqual(self._status_on(OWNER_A), 'CLEARED') + self.assertEqual( + self._status_on(OWNER_B), 'CLEARED', + 'the bulk clear stopped after one record of the shared code', + ) + + def test_07_the_openapi_document_declares_the_record_contract(self): + """The served spec says what the routes now do. + + @verifies REQ_INTEROP_013 + """ + spec = self.poll_endpoint_until('/docs', lambda d: d if 'openapi' in d else None) + + item = spec['components']['schemas']['FaultListItem'] + self.assertIn( + 'source_id', item['properties'], + 'the published fault item must declare the owner it carries', + ) + + route = spec['paths']['/apps/{app_id}/faults/{fault_code}'] + # The document has one response object per status, so a lock 409 and an + # ambiguous-fault 409 on the DELETE cannot be told apart. This gateway + # runs with locking off, which drops the lock contract from the document + # altogether. Ground that first: with it in place, a 409 on either + # method is there only because the route declares the ambiguous-fault + # refusal, and both checks below can fail. + root = self.get_json('') + self.assertIs( + root['capabilities']['locking'], False, + 'this launch must run with locking off, or the DELETE 409 below may be the lock one', + ) + self.assertNotIn( + 'x-medkit-lock-guarded', route['delete'], + 'with locking off the DELETE must carry no lock contract', + ) + for method in ('get', 'delete'): + self.assertIn( + '409', route[method]['responses'], + f'the per-entity fault {method} answers 409 on an ambiguous code ' + 'and has to declare it', + ) + + def test_075_the_detail_names_the_owner_and_not_the_lists_source_id(self): + """x-medkit.owner is the record's owner, source_id belongs to a list. + + @verifies REQ_INTEROP_013 + """ + self._raise_both() + + detail = self.get_json(f'/apps/{OWNER_A}/faults/{SHARED_CODE}') + + self.assertEqual(detail['x-medkit']['owner'], OWNER_A) + self.assertNotIn('source_id', detail['x-medkit']) + + # The list-level key is the other meaning, and it is still there. + listing = self.get_json(f'/apps/{OWNER_A}/faults') + self.assertIn('source_id', listing['x-medkit']) + + def test_076_every_recording_download_declares_its_409(self): + """The download routes answer 409 on an ambiguous code and say so. + + A GET carries no lock marker, so a 409 on these operations is there + only because the route declares the ambiguous-fault refusal. + + @verifies REQ_INTEROP_072 + """ + spec = self.poll_endpoint_until('/docs', lambda d: d if 'openapi' in d else None) + + downloads = { + path: item['get'] for path, item in spec['paths'].items() + if path.endswith('/bulk-data/{category_id}/{file_id}') and 'get' in item + } + # One per entity type that serves bulk data, so the loop below cannot + # pass by matching nothing. + self.assertEqual(len(downloads), 6, sorted(downloads)) + for path, operation in sorted(downloads.items()): + self.assertIn( + '409', operation['responses'], + f'GET {path} answers 409 x-medkit-ambiguous-fault and has to declare it', + ) + + def test_077_a_bare_code_bag_url_refuses_to_pick_an_owner(self): + """A recording URL carrying a bare fault code names no single record. + + Two owners of one code in the component's scope means the compatibility + URL (the pre-recording-id form, which carries a fault code) addresses + neither record, and each owner keeps its own recordings. Serving the + lowest-sorting owner's bytes under that URL would never say whose they + were, so it answers the same 409 the fault routes do. Rosbag capture is + off in this launch, so the assertion is on the refusal the resolution + makes before any bag is looked up. + + @verifies REQ_INTEROP_072 + """ + self._raise_both() + + response = requests.get( + f'{self.BASE_URL}/components/{HOST_COMPONENT}/bulk-data/rosbags/{SHARED_CODE}', + timeout=10, + ) + + self.assertEqual(response.status_code, 409, response.text) + body = response.json() + self.assertEqual(body.get('vendor_code'), 'x-medkit-ambiguous-fault', body) + self.assertEqual( + sorted(body['parameters']['owners']), [OWNER_A, OWNER_B], body) + + # One owner in scope is one record, so that entity's own URL is not + # ambiguous - it 404s only because no bag was captured. + single = requests.get( + f'{self.BASE_URL}/apps/{OWNER_A}/bulk-data/rosbags/{SHARED_CODE}', timeout=10 + ) + self.assertNotEqual( + single.status_code, 409, + 'one owner in scope must not read as ambiguous', + ) + + def test_08_stream_frames_name_the_record_they_describe(self): + """Each frame carries the entity hint of its own record's owner.""" + frames = [] + stop_event = threading.Event() + response = requests.get( + f'{self.BASE_URL}/faults/stream', stream=True, timeout=(5, 60) + ) + self.assertEqual(response.status_code, 200) + pump = threading.Thread( + target=self._pump_stream, args=(response, frames, stop_event), daemon=True + ) + pump.start() + try: + self._prime_stream(frames) + self._report(OWNER_A, fault_code=SHARED_CODE) + self._report(OWNER_B, fault_code=SHARED_CODE) + + seen = {} + deadline = time.monotonic() + FAULT_TIMEOUT + while time.monotonic() < deadline and len(seen) < 2: + for frame in list(frames): + data = frame.get('data') + if data is None: + continue + payload = json.loads(data) + fault = payload.get('fault', {}) + if fault.get('fault_code') != SHARED_CODE: + continue + hint = payload.get('x-medkit') + self.assertIsNotNone( + hint, f'a frame for {SHARED_CODE} carried no entity hint: {payload}' + ) + self.assertEqual( + hint['entity_id'], fault.get('source_id'), + 'the hint must name the record the frame describes', + ) + seen[hint['entity_id']] = hint + time.sleep(0.2) + + self.assertEqual( + sorted(seen), [OWNER_A, OWNER_B], + f'both owners must appear on the stream, saw: {sorted(seen)}', + ) + for hint in seen.values(): + self.assertEqual(hint['entity_type'], 'apps') + finally: + stop_event.set() + response.close() + pump.join(timeout=5) + + @staticmethod + def _pump_stream(response, frames, stop_event): + """Collect SSE frames as dicts of field -> value.""" + current = {} + try: + for line in response.iter_lines(decode_unicode=True): + if stop_event.is_set(): + break + if line is None: + continue + if line == '': + if current: + frames.append(current) + current = {} + continue + if line.startswith(':'): + continue # keepalive comment + key, _, value = line.partition(':') + current[key.strip()] = value.strip() + except Exception: # noqa: BLE001 - closed socket on test teardown + pass + + def _prime_stream(self, frames): + """Block until the fault event pipeline demonstrably reaches the stream. + + /fault_manager/events is reliable but volatile: an event published + before the gateway's subscription has matched the fault manager's + publisher is lost outright. Repeating a sacrificial record until one of + its frames arrives proves service -> fault manager -> events -> SSE. + """ + deadline = time.monotonic() + FAULT_TIMEOUT + while time.monotonic() < deadline: + self._report(OWNER_A, fault_code=PRIME_CODE, times=2) + settle = min(time.monotonic() + 1.0, deadline) + while time.monotonic() < settle: + for frame in list(frames): + data = frame.get('data') + if data is None: + continue + if json.loads(data).get('fault', {}).get('fault_code') == PRIME_CODE: + return + time.sleep(0.1) + raise AssertionError( + f'no event for priming fault {PRIME_CODE} on /faults/stream within ' + f'{FAULT_TIMEOUT}s, events pipeline never went live' + ) + + +@launch_testing.post_shutdown_test() +class TestProcessOutput(unittest.TestCase): + """All processes exited cleanly (SIGTERM allowed for SSE teardown).""" + + def test_exit_codes(self, proc_info): + for process_name in proc_info.process_names(): + self.assertIn( + proc_info[process_name].returncode, ALLOWED_EXIT_CODES, + f'{process_name} exited with {proc_info[process_name].returncode}', + ) diff --git a/src/ros2_medkit_integration_tests/test/features/test_grouping_entity_aggregation.test.py b/src/ros2_medkit_integration_tests/test/features/test_grouping_entity_aggregation.test.py index 9587e1171..e61399616 100644 --- a/src/ros2_medkit_integration_tests/test/features/test_grouping_entity_aggregation.test.py +++ b/src/ros2_medkit_integration_tests/test/features/test_grouping_entity_aggregation.test.py @@ -33,14 +33,16 @@ ADDRESSING SECOND, and it has exactly two cases, decided per collection rather than per handler: - a LEAF-OWNED item - a topic, an operation, a configuration key - belongs to - one leaf, and is addressed ":". - an AGGREGATE item - a fault, which has one code and a SET of reporting - sources - belongs to the aggregate itself, and is + a LEAF-OWNED item - a topic, an operation, a configuration key, a fault + record - belongs to one leaf, and is addressed + ":". A fault record is (fault_code, + reporting source), and the source is the leaf. + an AGGREGATE item - one a grouping computes over its members rather than + holds - belongs to the aggregate itself, and is addressed by its own id while naming its contributors. -Faults, logs and bulk-data are not exceptions to the model; they are the second -case. Which case a collection is, is declared once. +Logs and bulk-data are not exceptions to the model. Which case a collection is, +is declared once. WHAT THIS SUITE EXISTS TO PREVENT, all three measured on this topology: diff --git a/src/ros2_medkit_integration_tests/test/features/test_rosbag_owner_download.test.py b/src/ros2_medkit_integration_tests/test/features/test_rosbag_owner_download.test.py new file mode 100644 index 000000000..bce0f13e7 --- /dev/null +++ b/src/ros2_medkit_integration_tests/test/features/test_rosbag_owner_download.test.py @@ -0,0 +1,194 @@ +#!/usr/bin/env python3 +# Copyright 2026 bburda +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""A recording downloads for the owner of the record it is attached to. + +Two external apps under one component report the same fault code, each under +its own entity id, so the fault manager keeps two records of that code and +captures a recording for each. A recording belongs to the (fault code, owner) +records it is attached to, so over the recording-id URL: + +* each app downloads its own recording and gets ``404`` on the other's, in both + directions, although both apps own a record of the same code +* the component hosting both apps downloads both recordings +* after one app's record is cleared, that app still gets ``404`` on the other + app's recording, and still downloads its own + +The unit layer pins the same rule against a stub fault manager +(``PluginClearOwnerTest.EachOwnerDownloadsOnlyItsOwnRecordingById``). This runs +it against the real fault manager and its real rosbag rows. +""" + +import os +import time +import unittest + +from ament_index_python.packages import get_package_share_directory +import launch_testing +import rclpy +from rclpy.node import Node +import requests +from ros2_medkit_msgs.msg import Fault +from ros2_medkit_msgs.srv import ReportFault + +from ros2_medkit_test_utils.constants import ALLOWED_EXIT_CODES +from ros2_medkit_test_utils.gateway_test_case import GatewayTestCase +from ros2_medkit_test_utils.launch_helpers import create_test_launch + + +OWNER_A = 'owner-a' +OWNER_B = 'owner-b' +HOST_COMPONENT = 'shared-code-hub' +SHARED_CODE = 'SHARED_RECORDED' +RECORDING_TIMEOUT = 30.0 + + +def generate_test_description(): + manifest_path = os.path.join( + get_package_share_directory('ros2_medkit_gateway'), + 'config', 'examples', 'fault_owner_identity_manifest.yaml', + ) + return create_test_launch( + # The lidar publishes the scan topic the fault manager records, so each + # confirmation leaves a real bag behind. + demo_nodes=['lidar_sensor'], + fault_manager=True, + fault_manager_params={ + 'confirmation_threshold': -1, # a single report confirms + # Keep a cleared record's recording, so the app that cleared can + # still be shown downloading its own bag after the clear. + 'snapshots.rosbag.auto_cleanup': False, + }, + gateway_params={ + 'discovery.mode': 'hybrid', + 'discovery.manifest_path': manifest_path, + 'discovery.manifest_strict_validation': False, + }, + ) + + +class TestRosbagOwnerDownload(GatewayTestCase): + """Each owner of one fault code downloads only its own recording by id.""" + + MIN_EXPECTED_APPS = 2 + REQUIRED_APPS = {OWNER_A, OWNER_B} + + @classmethod + def setUpClass(cls): + rclpy.init() + cls._reporter = Node('rosbag_owner_download_reporter') + cls._report_client = cls._reporter.create_client( + ReportFault, '/fault_manager/report_fault' + ) + super().setUpClass() + assert cls._report_client.wait_for_service(timeout_sec=15.0), \ + 'report_fault service not available' + + @classmethod + def tearDownClass(cls): + cls._reporter.destroy_node() + rclpy.shutdown() + + def _report(self, source_id): + """Report SHARED_RECORDED under one owner, fire and forget.""" + req = ReportFault.Request() + req.fault_code = SHARED_CODE + req.event_type = ReportFault.Request.EVENT_FAILED + req.severity = Fault.SEVERITY_ERROR + req.description = f'reported by {source_id}' + req.source_id = source_id + self._report_client.call_async(req) + rclpy.spin_once(self._reporter, timeout_sec=0.1) + + def _recordings_listed_for(self, app_id): + listing = self.get_json(f'/apps/{app_id}/bulk-data/rosbags') + return sorted( + item['id'] for item in listing.get('items', []) + if SHARED_CODE in item.get('x-medkit', {}).get('fault_codes', []) + ) + + def _record_for(self, app_id): + """Raise the app's own record and wait for its one finished recording.""" + deadline = time.monotonic() + RECORDING_TIMEOUT + ids = [] + while time.monotonic() < deadline: + # Re-sent until the recording is listed: the report is fire and + # forget, and a repeat inside one occurrence changes nothing. + self._report(app_id) + ids = self._recordings_listed_for(app_id) + if ids: + break + time.sleep(0.5) + self.assertEqual(len(ids), 1, f'{app_id} should hold exactly one recording, got {ids}') + return ids[0] + + def test_01_each_owner_downloads_only_its_own_recording(self): + """Two owners of one code, a recording each, over the recording-id URL. + + @verifies REQ_INTEROP_072 + """ + # One after the other, so the second confirmation falls outside the + # first recording's post-roll and gets a recording of its own instead + # of attaching to the first one as a burst would. + rec_a = self._record_for(OWNER_A) + rec_b = self._record_for(OWNER_B) + self.assertNotEqual(rec_a, rec_b, 'the two records share one recording') + + self._assert_downloads({ + (f'/apps/{OWNER_A}', rec_a): 200, + (f'/apps/{OWNER_A}', rec_b): 404, + (f'/apps/{OWNER_B}', rec_b): 200, + (f'/apps/{OWNER_B}', rec_a): 404, + (f'/components/{HOST_COMPONENT}', rec_a): 200, + (f'/components/{HOST_COMPONENT}', rec_b): 200, + }) + + # Clearing owner-a's record keeps owner-a in its own scope with a + # CLEARED record of the code. That must not make owner-b's recording + # owner-a's, and owner-a's own recording stays served. + self.delete_request(f'/apps/{OWNER_A}/faults/{SHARED_CODE}', expected_status=204) + self._assert_downloads({ + (f'/apps/{OWNER_A}', rec_b): 404, + (f'/apps/{OWNER_A}', rec_a): 200, + (f'/apps/{OWNER_B}', rec_a): 404, + (f'/apps/{OWNER_B}', rec_b): 200, + }) + + def _assert_downloads(self, expected): + """Download each (entity path, recording id) and report every wrong status at once.""" + wrong = [] + for (entity_path, recording_id), want in expected.items(): + response = requests.get( + f'{self.BASE_URL}{entity_path}/bulk-data/rosbags/{recording_id}', + timeout=15, + ) + if response.status_code != want: + wrong.append( + f'{entity_path} downloading {recording_id} answered ' + f'{response.status_code}, expected {want}' + ) + self.assertEqual(wrong, [], '\n'.join(wrong)) + + +@launch_testing.post_shutdown_test() +class TestProcessOutput(unittest.TestCase): + """All processes exited cleanly.""" + + def test_exit_codes(self, proc_info): + for process_name in proc_info.process_names(): + self.assertIn( + proc_info[process_name].returncode, ALLOWED_EXIT_CODES, + f'{process_name} exited with {proc_info[process_name].returncode}', + ) diff --git a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/README.md b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/README.md index 117a403ff..f5dc8c2d0 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/README.md +++ b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/README.md @@ -217,13 +217,14 @@ directly - the obvious choice - makes them reachable from no endpoint at all: it. There is no server-level `/faults/{code}` route either, so the flat `/faults` list was the only place these faults existed. -Hanging each fault on the FQN of the node it is ABOUT does not work either: fault -identity in the store is `fault_code` alone (`fault_code TEXT PRIMARY KEY`) while -`reporting_sources` accumulates into a set on that one record, and `fault_in_source_scope` -requires EVERY source to be in scope - so two dead nodes would hide -`GRAPH_NODE_DISAPPEARED` from both of their `/apps//faults` pages. One owned entity -per code keeps exactly one reporting source per record; the affected nodes are named in -the description. +Hanging each fault on the FQN of the node it is ABOUT is a different design, not an +impossible one: a record is (`fault_code`, reporting source), so raising +`GRAPH_NODE_DISAPPEARED` under each dead node's FQN would give each of them its own +record on its own `/apps//faults` page. This detector aggregates by choice. A +graph-level condition is one condition however many nodes it names, the affected set +changes every tick, and per-node records would make the operator reconstruct the graph +view from a list raising and clearing under them. One owned entity per code keeps the +condition addressable as one thing, and the affected nodes are named in the description. ### Detectors @@ -298,9 +299,8 @@ means the subscriber does not constrain that policy, so it is always compatible. and depth are not RxO-compatibility dimensions and are deliberately not checked. **Aggregated fault, not per-topic**, via the shared `AggregatedFault` helper -(`aggregated_fault.hpp`): the fault_manager -identifies a fault by `fault_code` alone, so one `GRAPH_QOS_MISMATCH` per mismatched -topic would collide into a single record under the shared code. One graph-level fault +(`aggregated_fault.hpp`): a mismatch is a property of the graph, not of one topic, and +the mismatched set changes every tick. One graph-level fault enumerates every currently-mismatched topic; it clears (`EVENT_PASSED`) on every tick where nothing is mismatched, subject to the same `healing_enabled` requirement described in "Closing the loop" above to actually reach HEALED. @@ -415,8 +415,8 @@ plugins: ``` **Aggregated fault, not per-pair**, via the same `AggregatedFault` helper `qos_mismatch` -uses - the fault_manager identifies a fault by `fault_code` alone, so one `GRAPH_ORPHAN` -per near-miss pair would collide into a single record under the shared code. One +uses, and for the same reason: an orphaned pair is a property of the graph and the +orphaned set changes every tick. One graph-level fault enumerates every currently-orphaned pair, keyed by the canonical ` <-> ` string; it clears (`EVENT_PASSED`) on every tick where nothing is orphaned, subject to the same `healing_enabled` requirement @@ -473,8 +473,8 @@ whoever reads the fault. **Aggregated fault, not per-node.** All drift in the graph is one graph-level `GRAPH_PARAM_DRIFT`, not a fault per owning node, for the same reason as the other -detectors: the fault_manager identifies a fault by `fault_code` alone, so per-node faults -would collide into a single record. The description enumerates every currently-drifted +detectors: drift is a graph-level condition and the drifted set changes every tick. +The description enumerates every currently-drifted `(node, parameter)` pair, up to a cap of 480 characters - past that the text is truncated. Each app's own contribution is trimmed to 150 characters before that, so one node drifting on many parameters cannot fill the 480 by itself, and the entries are ordered `expect` violations first, @@ -1043,9 +1043,9 @@ the tick thread, the only thread that ever pumps these callbacks, so nothing it queued is delivered afterwards. **Aggregated fault, not per-node**, via the shared `AggregatedFault` helper - the same -rationale as every other detector here: the fault_manager identifies a fault by -`fault_code` alone, so one `GRAPH_NODE_INACTIVE` per stuck node would collide into a -single record under the shared code. Three graph-level faults, each enumerating every +rationale as every other detector here: a stuck node is reported as part of a +graph-level condition whose affected set changes every tick. +Three graph-level faults, each enumerating every currently affected node for ITS OWN code, each description capped independently at 480 characters with a truncation marker. Like `orphan` and `param_drift`, all three are fixed-severity `AggregatedFault` members at class scope (`aggregated_inactive_`, @@ -1065,7 +1065,7 @@ independently.** The description used to list affected nodes in fqn order interesting but not once the cap is full: a fleet sharing `require_active: ["controller_server"]` across a dozen robots fills the 480-char cap from the alphabetically-earliest ones, and a THIRTEENTH robot going inactive afterward would -be silently invisible forever - one shared `fault_code`, one record, no way to tell the +be silently invisible forever - one aggregated fault, one description, no way to tell the operator which of thirteen actually broke. The tracker reports which fqns entered EACH fault's content on THIS tick - `newly_affected`, `newly_unreadable`, `newly_not_managed` - and each list orders only its OWN fault's `AggregatedFault::emit_ordered` call, since the @@ -1526,10 +1526,11 @@ SAME occurrence continuing, not a new one, so restarting the dead node fast enou before anyone acknowledges it will not move the count. The honest limits, read off this branch's fault-manager storage rather than assumed: the -per-fault rosbag store enforces `fault_code` as UNIQUE, so a fault can hold at most one -recording at a time - a later confirmation's capture replaces the earlier one on disk rather -than accumulating a history. The freeze frame is the same shape: one row per `fault_code`, -overwritten on every capture, so a fifth occurrence's captured values overwrite the first's. +rosbag store is unique on `(fault_code, owner, file_path)`, so one record keeps a history of +recordings and the cap governs how many. The freeze frame is one row per record +(`(fault_code, owner)`), overwritten on every capture, so a fifth occurrence's captured +values overwrite the first's. Both are per record, so another source reporting the same +code keeps its own. What survives every occurrence by default is the count itself; a hash-chained record of every raise/clear/heal transition also exists, but only once the fault manager's own `audit_log.enabled` is turned on, which it is not by default. Recordings and the freeze @@ -1563,10 +1564,9 @@ operator, not by the passage of a restart. Re-seeding the tracker from the fault store at startup would change it, and today it cannot be done. `/fault_manager/list_faults` would tell this detector that a `GRAPH_NODE_DISAPPEARED` -record is outstanding and that it is among its reporting sources - but not WHICH nodes it -names, which is the only thing that would make a fresh instance's silence meaningful. The fault -manager keeps one record per `fault_code`, and this detector folds every dead key into that one -record's description. That description is this detector's own deterministic text, so for a +record is outstanding and that it owns it - but not WHICH nodes it names, which is the only +thing that would make a fresh instance's silence meaningful. This detector raises one +aggregated fault under one owner and folds every dead key into that record's description. That description is this detector's own deterministic text, so for a record that never hit `kMaxDescriptionChars` the key list could in principle be read back out of it - the detector does not, because past the cap the remainder is collapsed into a count and the names are gone for good, and a re-seeding rule that works only for small faults is worse than diff --git a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/design/graph_watchdog.rst b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/design/graph_watchdog.rst index bba329b45..c7f8cb0fa 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/design/graph_watchdog.rst +++ b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/design/graph_watchdog.rst @@ -13,7 +13,7 @@ below). Structure --------- - **Plugin shell** (``GraphWatchdogPlugin``): loads via the gateway plugin ABI - (v7). In ``set_context`` it casts the context with ``as_ros_plugin_context``, + (v8). In ``set_context`` it casts the context with ``as_ros_plugin_context``, creates one ``rclcpp::Client`` on the gateway node, and starts a dedicated tick thread. The tick is deliberately NOT a gateway wall timer: detectors do blocking parameter/service reads, and the gateway's small @@ -194,9 +194,9 @@ reported by a live endpoint, which always carries the resolved profile) never ra **Aggregation via the shared helper.** ``qos_mismatch`` was the first detector to use the ``AggregatedFault`` helper (``aggregated_fault.hpp``); ``orphan`` and ``param_drift`` go through the same one, so no detector reimplements the level-triggered raise/clear -pattern. The rationale: the fault_manager identifies a fault by -``fault_code`` alone, so one ``GRAPH_QOS_MISMATCH`` per mismatched topic would collide -into a single record under the shared code. The detector keeps one ``AggregatedFault`` +pattern. The rationale: a mismatch is a property of the graph, not of one topic, and +the mismatched set changes every tick, so a fault per topic would raise and clear +under the operator rather than describe a condition. The detector keeps one ``AggregatedFault`` instance for the whole graph and, each tick, hands it every currently-mismatched topic's description (keyed by topic name so a repeat mismatch on the same topic overwrites rather than duplicates); an empty map on a clean tick clears @@ -682,7 +682,7 @@ log wants to know WHICH of the two is happening. **Three independent faults, not one shared record.** ``GRAPH_NODE_INACTIVE``, ``GRAPH_NODE_UNREADABLE`` and ``GRAPH_NODE_NOT_MANAGED`` are each raised through the shared ``AggregatedFault`` helper every ``GRAPH_*`` detector uses (one graph-level -record per code, since the fault_manager identifies a fault by ``fault_code`` alone), +record per code, because the condition is graph-level and its affected set moves), but as three SEPARATE, fixed-severity class members - ``GRAPH_NODE_INACTIVE`` always ``SEVERITY_ERROR``, the other two always ``SEVERITY_WARN`` - the same shape ``orphan_detector`` and ``param_drift_detector`` use, rather than one record whose @@ -730,7 +730,7 @@ independently.** The description used to list affected nodes in fqn order interesting but not once the cap is full: a fleet sharing ``require_active: ["controller_server"]`` across a dozen robots fills the 480-char cap from the alphabetically-earliest ones, and a THIRTEENTH robot going inactive afterward -would be silently invisible forever - one shared ``fault_code``, one record, no way to +would be silently invisible forever - one aggregated fault, one description, no way to tell the operator which of thirteen actually broke. The tracker reports which fqns entered EACH fault's content on THIS tick - ``newly_affected``, ``newly_unreadable``, ``newly_not_managed`` - and each list orders only its OWN fault's diff --git a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/include/ros2_medkit_graph_watchdog/aggregated_fault.hpp b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/include/ros2_medkit_graph_watchdog/aggregated_fault.hpp index a1fc860fc..d7970eb28 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/include/ros2_medkit_graph_watchdog/aggregated_fault.hpp +++ b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/include/ros2_medkit_graph_watchdog/aggregated_fault.hpp @@ -37,12 +37,14 @@ namespace ros2_medkit_graph_watchdog { /// `/faults/{code}` route either, so the flat `/faults` list was the only place these /// faults existed. /// -/// Hanging each fault on the FQN of the node it is ABOUT does not work either: -/// fault identity in the store is `fault_code` alone (`fault_code TEXT PRIMARY KEY`), -/// and `reporting_sources` accumulates into a set on that one record while -/// `fault_in_source_scope` requires EVERY source to be in scope. Two dead nodes would -/// therefore hide GRAPH_NODE_DISAPPEARED from both of their `/apps//faults` pages. -/// One owned entity per code keeps exactly one source per record. +/// Hanging each fault on the FQN of the node it is ABOUT is a different design, +/// not an impossible one: a record is (fault_code, reporting source), so raising +/// GRAPH_NODE_DISAPPEARED under each dead node's FQN would give each of them its +/// own record on its own `/apps//faults` page. This detector aggregates by +/// choice. A graph-level condition is one condition however many nodes it names, +/// the affected set changes every tick, and per-node records would make the +/// operator reconstruct the graph view from a list that raises and clears under +/// them. One owned entity per code keeps the condition addressable as one thing. /// /// Same shape as the ADS plugin's device entity, for the same reason. inline constexpr const char * kGraphWatchdogEntityId = "graph_watchdog"; diff --git a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/test/e2e/test_node_death_e2e.test.py b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/test/e2e/test_node_death_e2e.test.py index 5415c039d..9f4d2ae9b 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/test/e2e/test_node_death_e2e.test.py +++ b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/test/e2e/test_node_death_e2e.test.py @@ -1313,15 +1313,12 @@ class TestNodeDeathRestartLoopOccurrences(unittest.TestCase): occurrence"), which is why this scenario acknowledges explicitly between cycles instead of waiting for an organic heal. - Deliberately does NOT assert the "more than one recording" half of this row. The - per-fault rosbag store enforces `fault_code` as UNIQUE (`sqlite_fault_storage.cpp`'s - `store_rosbag_file_locked`: "INSERT OR REPLACE INTO rosbag_files ... (fault_code is - UNIQUE)" - the SAME fault_code can hold at most one recording ROW, structurally, and a - re-confirm deletes the previous bag file from disk). No config key in the fault manager - lifts this cap today: recording more than one rosbag per fault_code needs the storage - schema itself to change. Asserting a recording count here would either assert something - trivially true for the wrong reason (no captures happen at all without a detector) or - something structurally impossible to ever pass - neither is written. + Deliberately does NOT assert the "more than one recording" half of this row. How many + recordings one record keeps is governed by the fault manager's own retention + configuration, not by anything this scenario drives, so asserting a count here would + assert the retention default rather than the detector. What this scenario is about is + the occurrence count, which is what the acknowledge-between-cycles sequence above + establishes. """ def test_repeated_kills_reach_matching_occurrence_count(self, target_node): diff --git a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/test/test_node_death_integration.cpp b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/test/test_node_death_integration.cpp index d6098ead2..3a7ea3639 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/test/test_node_death_integration.cpp +++ b/src/ros2_medkit_plugins/ros2_medkit_graph_watchdog/test/test_node_death_integration.cpp @@ -1171,9 +1171,9 @@ TEST_F(NodeDeathIntegrationTest, UngatedClearStaysSuppressedWhileAnUnrelatedNode // Why it is not simply fixed here: the guard exists because a fresh instance cannot tell "the // node came back" from "I never saw that node". Answering that needs the KEYS the outstanding // record names, and they are not reachable. /fault_manager/list_faults would tell this detector -// that a GRAPH_NODE_DISAPPEARED record is outstanding and that it is among its reporting -// sources, but not which nodes it names: fault_manager aggregates one record per fault_code, -// and this detector merges every dead key into that one record's description, capped at +// that a GRAPH_NODE_DISAPPEARED record is outstanding and that this plugin owns it, but not +// which nodes it names: this detector raises one aggregated fault under one owner and merges +// every dead key into that record's description, capped at // AggregatedFault::kMaxDescriptionChars with the remainder collapsed into a count. Acting on // the record's mere existence would clear it for a node that is still dead and simply never // re-observed, which is the ungated heal the sibling tests above forbid. diff --git a/src/ros2_medkit_plugins/ros2_medkit_opcua/README.md b/src/ros2_medkit_plugins/ros2_medkit_opcua/README.md index 748d8da48..054a9157a 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_opcua/README.md +++ b/src/ros2_medkit_plugins/ros2_medkit_opcua/README.md @@ -192,7 +192,7 @@ A write-capable build restores everything above; nothing else differs. | GET | `/apps/{id}/x-plc-data` | All OPC-UA values for entity (with units, types, timestamps) | | GET | `/apps/{id}/x-plc-data/{name}` | Single data point value | | POST | `/apps/{id}/x-plc-operations/set_{name}` | Write value to PLC (`{"value": 75.0}`) - write-capable build only; not registered otherwise | -| GET | `/components/{id}/x-plc-status` | Connection state, poll stats, active alarms, and `write_capable` - the write surface of the plugin object itself | +| GET | `/components/{id}/x-plc-status` | Connection state, poll stats, active alarms, and `write_capable` - the write surface of the plugin object itself. Served for the plugin's own component only. Any other component answers 404 `resource-not-found` | ### Standard SOVD (provided by gateway) @@ -255,6 +255,12 @@ build, `true` only in one built with `-DMEDKIT_OPCUA_READ_ONLY=OFF`. An absent `x-plc-operations` capability alone does not say this, because a write-capable build whose node map marks nothing writable shows the same absence. +The status describes the plugin's own OPC UA session, so it is served only under +the component the plugin introspects (`openplc_runtime` above). Any other +component in the gateway, for example one another plugin introspects, answers +404 `resource-not-found`, the same answer the data route gives an entity with no +mapped points. + ## Finding Node IDs on your PLC The plugin identifies PLC tags by OPC-UA node IDs in the canonical string @@ -375,13 +381,14 @@ nodes: # config, so the loader logs a warning and skips that bit rule while still loading # the rest of the config. # -# Fault codes must be globally unique across ALL fault sources - every `alarm` / -# `status_bits` / `fault_enum` entry AND every `event_alarms` entry (including -# each of its `mappings[].fault_code`) - regardless of which entity owns them. The fault manager keys and clears faults by -# fault_code alone, so a code reused on two sources (even on different entities, -# even one polled and one event-driven) would flap raise/clear or clear the other -# source's fault. The loader rejects the whole file at load with an actionable -# error naming both sources. +# Fault codes must be unique across ALL fault sources in this file - every +# `alarm` / `status_bits` / `fault_enum` entry AND every `event_alarms` entry +# (including each of its `mappings[].fault_code`) - regardless of which entity +# owns them. The reason is this plugin's own shared `FaultTransitionTracker`, +# which is keyed by fault_code alone: two sources emitting one code would +# alternately raise and clear it every cycle, before either report is sent. The +# loader rejects the whole file at load with an actionable error naming both +# sources. # Native OPC-UA AlarmConditionType events (issue #386). Subscribes to alarms # defined inside the PLC (Siemens Program_Alarm / ProDiag, Beckhoff TF6100, @@ -585,9 +592,9 @@ no node-map file required. SourceNode + EventType + Message. The SourceNode is folded into every tier so two distinct conditions sharing a ConditionName/SourceName but sitting on different sources (e.g. two identical FB instances each raising - "Overpressure") never collapse onto one code - the fault manager keys/clears - by code alone, so a collision would let one condition's clear wipe the - other's still-active fault. Message is NOT folded into the slug tiers, so a + "Overpressure") never collapse onto one code - the shared transition tracker + is keyed by code alone, so a collision would flap raise/clear between the two + conditions every cycle. Message is NOT folded into the slug tiers, so a condition whose Message differs between its active and inactive notifications still maps to one code. The hash tier additionally folds in the Message deliberately - a real Siemens S7-1500 multiplexes every `Program_Alarm` of @@ -1004,10 +1011,12 @@ GET /api/v1/apps/tank_process/faults ``` When the value returns below threshold, the plugin calls the fault manager's -`~/clear_fault` service (`/fault_manager/clear_fault`) for that fault code. That -is the same service an operator's -`DELETE /api/v1/apps/{app_id}/faults/{fault_code}` ends up calling, so a device -de-assert is a clear like any other. It drops the fault's value snapshots unless +`~/clear_fault` service (`/fault_manager/clear_fault`) for that fault code and +the entity that reported it, which together address one record. That is the same +service an operator's `DELETE /api/v1/apps/{app_id}/faults/{fault_code}` ends up +calling, so a device de-assert is a clear like any other. That REST route reaches the plugin +through `FaultProvider::clear_fault_record` with the owner the gateway resolved, so on a +component route the clear names the hosted app that reported the record. It drops the fault's value snapshots unless `snapshots.retain_on_clear` is set, and its rosbag recording unless `snapshots.rosbag.auto_cleanup` is off or `snapshots.rosbag.max_bags_per_fault` keeps a history. diff --git a/src/ros2_medkit_plugins/ros2_medkit_opcua/design/index.rst b/src/ros2_medkit_plugins/ros2_medkit_opcua/design/index.rst index b7da7bac0..3c49e0be6 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_opcua/design/index.rst +++ b/src/ros2_medkit_plugins/ros2_medkit_opcua/design/index.rst @@ -92,7 +92,13 @@ edge-triggered callbacks: - **Active -> fault reported** via ``/fault_manager/report_fault`` with the configured severity and message -- **Cleared -> fault cleared** via ``/fault_manager/clear_fault`` by fault code +- **Cleared -> fault cleared** via ``/fault_manager/clear_fault`` by fault code and the + entity the plugin reported it under, which together name one record + +A REST ``DELETE /{entity}/faults/{code}`` on an entity this plugin owns reaches +``FaultProvider::clear_fault_record`` with the owner the gateway resolved, and the clear +names that owner. For a component route it is one of the component's hosted apps, not the +component itself. The two-argument ``clear_fault`` falls back to the addressed entity. The plugin keeps per-fault state only long enough to detect edges; the fault manager owns persistence and fault lifecycle. diff --git a/src/ros2_medkit_plugins/ros2_medkit_opcua/include/ros2_medkit_opcua/opcua_plugin.hpp b/src/ros2_medkit_plugins/ros2_medkit_opcua/include/ros2_medkit_opcua/opcua_plugin.hpp index a547740e8..5b23ad99f 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_opcua/include/ros2_medkit_opcua/opcua_plugin.hpp +++ b/src/ros2_medkit_plugins/ros2_medkit_opcua/include/ros2_medkit_opcua/opcua_plugin.hpp @@ -136,6 +136,8 @@ class OpcuaPlugin : public ros2_medkit_gateway::GatewayPlugin, const std::string & fault_code) override; tl::expected clear_fault(const std::string & entity_id, const std::string & fault_code) override; + tl::expected + clear_fault_record(const std::string & entity_id, const std::string & fault_code, const std::string & owner) override; // Resolve the SOVD severity bucket for an event alarm. An explicit configured // override wins; with none configured the raw OPC-UA event Severity (1-1000) @@ -177,7 +179,11 @@ class OpcuaPlugin : public ros2_medkit_gateway::GatewayPlugin, // Report/clear fault via ROS 2 service (private helpers, not the FaultProvider overrides) void send_report_fault(const std::string & entity_id, const std::string & fault_code, const std::string & severity_str, const std::string & message); - void send_clear_fault(const std::string & fault_code); + /// Clear one record, addressed by fault_code and the reporting source that + /// owns it. On the polled and event paths the owner is the entity this plugin + /// reported under. On the REST path it is the owner the gateway resolved, + /// which for a component route is one of its hosted apps. + void send_clear_fault(const std::string & owner, const std::string & fault_code); // Dispatch now if the fault_manager service is matched, else buffer the // dispatch (bounded, order-preserving) to be flushed once it appears. diff --git a/src/ros2_medkit_plugins/ros2_medkit_opcua/src/node_map.cpp b/src/ros2_medkit_plugins/ros2_medkit_opcua/src/node_map.cpp index 2f02eb216..c9b455657 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_opcua/src/node_map.cpp +++ b/src/ros2_medkit_plugins/ros2_medkit_opcua/src/node_map.cpp @@ -947,24 +947,24 @@ bool NodeMap::load(const std::string & yaml_path) { } } - // Global fault-code uniqueness validation. fault_manager keys and clears - // faults by ``fault_code`` ALONE (clear_fault(fault_code) / - // get_fault(fault_code)); the poller shares one ``FaultTransitionTracker`` - // that is likewise keyed by code alone. So a fault_code is a global - // identifier and must be unique across EVERY source that can emit it, - // regardless of entity_id: + // Global fault-code uniqueness validation. The reason is THIS PLUGIN's own + // ``FaultTransitionTracker``, which the poller shares across every node-map + // entry and which is keyed by ``fault_code`` alone: two entries emitting + // one code would alternately raise and clear it every cycle, because the + // tracker cannot tell their edges apart. So a fault_code must be unique + // across EVERY source this file declares, regardless of entity_id: // // * every polled detection fault (threshold + each status_bits bit + // each fault_enum code + the enum catch-all), and // * every native ``event_alarms`` subscription (issue #386). // - // Two sources sharing a code - even a polled code on entity_a and an - // event_alarms code on entity_b - would collide at fault_manager: one - // source's clear wipes the other's fault, or the two flap raise/clear - // every cycle. The (entity_id, fault_code) pair the earlier check keyed on - // is NOT sufficient, because the fault manager never sees entity_id in its - // key. Reject the whole file at load with an actionable error so the intent - // - one code, one source - is enforced before anything runs. + // The fault manager is not the reason: it keys a record by + // (fault_code, source_id), and this plugin reports and clears under the + // owning entity_id, so two entries sharing a code are two records there and + // neither clear touches the other. The shared tracker is what a collision + // breaks, and it breaks before either report is sent. Reject the whole file + // at load with an actionable error so the intent - one code, one source - + // is enforced before anything runs. { namespace fd = ros2_medkit::fault_detection; struct EmittedFault { @@ -977,9 +977,9 @@ bool NodeMap::load(const std::string & yaml_path) { if (!inserted) { RCLCPP_ERROR(rclcpp::get_logger("opcua.node_map"), "fault_code '%s' is emitted by more than one source (%s on entity '%s' and %s on " - "entity '%s'); fault codes must be globally unique across all detection entries and " - "event_alarms because the shared fault manager is keyed by code alone (a collision " - "clears the other source's fault or flaps raise/clear) - rename one of them", + "entity '%s'). Fault codes must be unique across all detection entries and " + "event_alarms in this file because the shared fault-transition tracker is keyed by " + "code alone (a collision flaps raise/clear every cycle). Rename one of them.", code.c_str(), it->second.pipeline, it->second.entity_id.c_str(), pipeline, entity_id.c_str()); return false; } diff --git a/src/ros2_medkit_plugins/ros2_medkit_opcua/src/opcua_plugin.cpp b/src/ros2_medkit_plugins/ros2_medkit_opcua/src/opcua_plugin.cpp index 3dff059a6..60fcb138b 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_opcua/src/opcua_plugin.cpp +++ b/src/ros2_medkit_plugins/ros2_medkit_opcua/src/opcua_plugin.cpp @@ -1018,6 +1018,15 @@ void OpcuaPlugin::handle_plc_status(const PluginRequest & req, PluginResponse & } std::shared_lock node_map_lock(node_map_mutex_); + // Everything below describes this plugin's own OPC UA session, which belongs + // to the one component the plugin introspects. Another component in the same + // entity tree exists but has no session here, so it gets the same 404 the + // data route gives an entity with no mapped points. + if (component_id != node_map_.component_id()) { + res.send_error(404, ERR_RESOURCE_NOT_FOUND, "No PLC status for component: " + component_id); + return; + } + auto snap = poller_->snapshot(); nlohmann::json j; @@ -1062,7 +1071,7 @@ void OpcuaPlugin::on_alarm_change(const std::string & entity_id, send_report_fault(entity_id, signal.fault_code, signal.severity, signal.message); } else { log_info("Alarm cleared: " + signal.fault_code + " on " + entity_id); - send_clear_fault(signal.fault_code); + send_clear_fault(entity_id, signal.fault_code); } } @@ -1240,8 +1249,8 @@ void OpcuaPlugin::on_event_alarm(const AlarmEventDelivery & delivery) { log_info("AlarmCondition HEALED (latched, awaiting ack/confirm): " + delivery.fault_code); break; case AlarmAction::ClearFault: - log_info("AlarmCondition CLEARED: " + delivery.fault_code); - send_clear_fault(delivery.fault_code); + log_info("AlarmCondition CLEARED: " + delivery.fault_code + " on " + delivery.entity_id); + send_clear_fault(delivery.entity_id, delivery.fault_code); break; case AlarmAction::NoOp: break; @@ -1281,17 +1290,57 @@ void OpcuaPlugin::send_report_fault(const std::string & entity_id, const std::st }); } -void OpcuaPlugin::send_clear_fault(const std::string & fault_code) { +void OpcuaPlugin::send_clear_fault(const std::string & owner, const std::string & fault_code) { if (!fault_clients_->clear) { log_warn("ClearFault service client not available"); return; } + // The record is (fault_code, source_id), and `owner` is that source. On the + // polled and event paths it is the entity this plugin reported under. On the + // REST path it is the owner the gateway resolved, which for a component route + // is one of the component's hosted apps and NOT the component itself. + // Clearing by code alone would instead reach whichever source the store + // resolved, which on a box running two plugin instances is another device's + // still-active fault. auto request = std::make_shared(); request->fault_code = fault_code; + request->source_id = owner; - send_or_buffer([this, request]() { - fault_clients_->clear->async_send_request(request); + send_or_buffer([this, request, owner, fault_code]() { + // Read the reply. The dispatch is still fire-and-forget as far as the + // caller is concerned, but a refusal has to reach the log: a clear the + // fault manager declined (no such record for that owner) used to leave no + // trace at all while the REST route answered as though it had worked. + // + // The reply callback holds nothing of the plugin. It runs on an executor + // thread whenever the reply arrives, which can be while the plugin is + // being destroyed, so it logs through a copy of the log sink taken here + // and carries the owner and the code by value. + auto warn = [sink = log_sink()](const std::string & msg) { + if (sink) { + sink(PluginLogLevel::kWarn, msg); + } + }; + using ClearFuture = rclcpp::Client::SharedFuture; + // async_send_request only accepts a callback whose parameter is exactly + // SharedFuture, by value (rclcpp checks the argument types, and a const + // reference is a different type), so the future cannot be taken by + // reference here. + // NOLINTNEXTLINE(performance-unnecessary-value-param) + auto on_reply = [warn = std::move(warn), owner, fault_code](ClearFuture future) { + try { + const auto & response = future.get(); + if (!response->success) { + warn("ClearFault refused for '" + fault_code + "' of source '" + owner + "': " + response->message); + } + } catch (const std::exception & e) { + warn("ClearFault reply for '" + fault_code + "' of source '" + owner + "' failed: " + e.what()); + } catch (...) { + warn("ClearFault reply for '" + fault_code + "' of source '" + owner + "' failed"); + } + }; + fault_clients_->clear->async_send_request(request, std::move(on_reply)); }); } @@ -1670,7 +1719,16 @@ tl::expected OpcuaPlugin::list_data( items.push_back(std::move(item)); } - return dto::DataListResult{nlohmann::json{{"items", std::move(items)}}}; + // Link state travels with the values, the same envelope the plugin's own data + // route serves. Without it a reader cannot tell a live value from the frozen + // last-known one the poller keeps across an outage: the fault-trigger engine + // holds a rule's state on a down link, and it can only do that if the content + // says the link is down. + nlohmann::json envelope{{"items", std::move(items)}}; + envelope["connected"] = snap.connected; + envelope["timestamp"] = std::chrono::system_clock::to_time_t(snap.timestamp); + + return dto::DataListResult{std::move(envelope)}; } tl::expected OpcuaPlugin::read_data(const std::string & entity_id, @@ -2059,11 +2117,10 @@ tl::expected OpcuaPlugin::list_fau item["severity"] = f.value("severity", 0); item["description"] = f.value("description", ""); item["status"] = f.value("status", ""); - // Fault records carry reporting_sources, not source_id. On this - // entity-scoped list prefer the requested entity when it is among the - // sources - reporting_sources[0] is ordering-dependent and can name a - // different co-reporting entity; fall back to the first source only - // when the entity itself never reported. + // A record carries its owner as source_id, which is also the single + // entry of reporting_sources. Only a record read without source_id falls + // back to reporting_sources: the requested entity when it is listed + // there, else the first entry. std::string source_id = f.value("source_id", ""); if (source_id.empty() && f.contains("reporting_sources") && f["reporting_sources"].is_array()) { for (const auto & src : f["reporting_sources"]) { @@ -2100,8 +2157,16 @@ tl::expected OpcuaPlugin::get_fa return tl::make_unexpected(FaultProviderErrorInfo{FaultProviderError::Internal, "plugin not initialized", 503}); } - // list_entity_faults returns a bare JSON array of fault objects scoped to - // this entity (see PluginContext contract). + // The entity match is the LIST, not a comparison here. list_entity_faults + // returns a bare JSON array already scoped to this entity (see PluginContext + // contract): the fault manager's records filtered by the entity's resolved + // source set, plus the peers' answers for this same entity. So an item of + // this code in that array is a record this entity owns, and another entity's + // record of the same code never reaches this loop. + // + // Comparing the item's source_id to entity_id instead would be wrong, not + // merely redundant: a COMPONENT owns the records its hosted apps reported, + // whose owner is the app id, and the comparison would 404 them. auto faults = ctx_->list_entity_faults(entity_id); if (faults.is_array()) { for (const auto & f : faults) { @@ -2111,17 +2176,31 @@ tl::expected OpcuaPlugin::get_fa } } - return tl::make_unexpected( - FaultProviderErrorInfo{FaultProviderError::FaultNotFound, "Fault not found: " + fault_code, 404}); + return tl::make_unexpected(FaultProviderErrorInfo{FaultProviderError::FaultNotFound, + "Fault not found: " + fault_code + " on entity " + entity_id, 404}); } tl::expected OpcuaPlugin::clear_fault(const std::string & entity_id, const std::string & fault_code) { + // The code-only form of the contract names no owner, so the addressed entity + // stands in, which is the source this plugin reports its own faults under. + return clear_fault_record(entity_id, fault_code, ""); +} + +tl::expected +OpcuaPlugin::clear_fault_record(const std::string & entity_id, const std::string & fault_code, + const std::string & owner) { if (!ctx_ || !fault_clients_) { return tl::make_unexpected(FaultProviderErrorInfo{FaultProviderError::Internal, "plugin not initialized", 503}); } - send_clear_fault(fault_code); + // The gateway resolved which record this route addresses, so its owner is + // what the clear names. The entity in the URL is not it: a component owns the + // records its hosted apps reported, and sending the component id addresses a + // record no source owns, which the fault manager declines. Only when the + // gateway resolved nothing (a record it does not hold) does the addressed + // entity stand in, which is what this plugin's own polled clears use. + send_clear_fault(owner.empty() ? entity_id : owner, fault_code); return dto::FaultClearResult{ nlohmann::json{{"status", "cleared"}, {"fault_code", fault_code}, {"entity_id", entity_id}}}; } diff --git a/src/ros2_medkit_plugins/ros2_medkit_opcua/src/opcua_poller.cpp b/src/ros2_medkit_plugins/ros2_medkit_opcua/src/opcua_poller.cpp index c24f9ca0b..515408d86 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_opcua/src/opcua_poller.cpp +++ b/src/ros2_medkit_plugins/ros2_medkit_opcua/src/opcua_poller.cpp @@ -685,13 +685,14 @@ bool OpcuaPoller::should_clear_after_refresh(SovdAlarmStatus last_status, const } // True when ANOTHER tracked condition instance carries the same fault_code and -// is currently Confirmed. Siemens A&C mints a fresh ConditionId (GUID) for -// every activation of the same Program_Alarm, so a stale instance's heal / -// reconcile-clear races the new instance's raise; because fault_manager keys -// by fault_code alone, letting the stale clear through would wipe the fault -// the live instance just asserted (seen on a real 1505SP: CONFIRMED then -// HEALED 170 us apart, alarm ends invisible). Caller must hold -// conditions_mutex_. +// is currently Confirmed. A server may mint a fresh ConditionId (GUID) for +// every activation of the same alarm, so a stale instance's heal or +// reconcile-clear races the new instance's raise. Both instances report under +// the entity that owns the code (the node map keeps every code on one entity), +// so they address one fault record, and letting the stale clear through would +// clear the record the live instance just raised (seen on a real controller: +// CONFIRMED then HEALED 170 us apart, the alarm ending invisible). Caller must +// hold conditions_mutex_. bool OpcuaPoller::same_code_active_elsewhere_locked(const std::string & fault_code, const std::string & condition_id_str) const { for (const auto & [cid, runtime] : conditions_) { diff --git a/src/ros2_medkit_plugins/ros2_medkit_opcua/test/test_node_map.cpp b/src/ros2_medkit_plugins/ros2_medkit_opcua/test/test_node_map.cpp index 2d2389844..523f95a68 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_opcua/test/test_node_map.cpp +++ b/src/ros2_medkit_plugins/ros2_medkit_opcua/test/test_node_map.cpp @@ -999,9 +999,11 @@ component_id: test TEST_F(NodeMapTest, RejectsCodeAcrossPipelinesDifferentEntities) { // Global-by-code uniqueness (issue #481): a polled detection code on // entity_a and an event_alarms code on entity_b share the SAME fault_code. - // fault_manager keys and clears by code alone, so the two collide even - // though the entities differ; the (entity_id, fault_code) pair the earlier - // guard keyed on let this through. The loader must reject the whole file. + // The poller's shared fault-transition tracker keys by code alone, so the + // two collide there even though the entities differ (the fault manager + // would keep them as two records). The (entity_id, fault_code) pair the + // earlier guard keyed on let this through. The loader must reject the whole + // file. std::string path = "/tmp/test_node_map_cross_entity_pipeline.yaml"; std::ofstream f(path); f << R"( diff --git a/src/ros2_medkit_plugins/ros2_medkit_opcua/test/test_opcua_plugin.cpp b/src/ros2_medkit_plugins/ros2_medkit_opcua/test/test_opcua_plugin.cpp index 0c6b58b64..9a49d9180 100644 --- a/src/ros2_medkit_plugins/ros2_medkit_opcua/test/test_opcua_plugin.cpp +++ b/src/ros2_medkit_plugins/ros2_medkit_opcua/test/test_opcua_plugin.cpp @@ -27,11 +27,13 @@ #include #include #include +#include #include #include #include #include #include +#include #include #include @@ -45,9 +47,24 @@ namespace ros2_medkit_gateway { +// What the route tests hand the stubs below. The provider tests never build a +// request or a response, so a null impl answers empty and records nothing. +struct StubRequest { + std::vector path_params; // index 0 is the full match +}; +struct StubResponse { + int status = 0; + std::string error_code; + nlohmann::json body; +}; + PluginRequest::PluginRequest(const void * impl) : impl_(impl) { } -std::string PluginRequest::path_param(size_t) const { +std::string PluginRequest::path_param(size_t index) const { + const auto * req = static_cast(impl_); + if (req != nullptr && index < req->path_params.size()) { + return req->path_params[index]; + } return {}; } std::string PluginRequest::header(const std::string &) const { @@ -67,9 +84,21 @@ std::string PluginRequest::query_param(const std::string &) const { PluginResponse::PluginResponse(void * impl) : impl_(impl) { } -void PluginResponse::send_json(const nlohmann::json &) { +void PluginResponse::send_json(const nlohmann::json & data) { + auto * res = static_cast(impl_); + if (res != nullptr) { + res->status = 200; + res->body = data; + } } -void PluginResponse::send_error(int, const std::string &, const std::string &, const nlohmann::json &) { +void PluginResponse::send_error(int status, const std::string & error_code, const std::string & message, + const nlohmann::json & /*parameters*/) { + auto * res = static_cast(impl_); + if (res != nullptr) { + res->status = status; + res->error_code = error_code; + res->body = {{"message", message}}; + } } // -- FakePluginContext -- @@ -217,6 +246,22 @@ component_id: test_runtime // -- DataProvider tests -- +// The provider's list envelope carries link state, the same shape the plugin's +// own data route serves. Without it a reader cannot tell a live value from the +// frozen last-known one the poller keeps across an outage, and the fault-trigger +// engine's link-down guard (content_reports_disconnected) never fires for a +// provider-served entity: a threshold rule then evaluates on a stale number for +// the whole outage. The fixture endpoint never connects, so the poller reports +// disconnected here. +TEST_F(OpcuaPluginTest, ListDataEnvelopeCarriesLinkState) { + auto result = plugin_.list_data("tank"); + + ASSERT_TRUE(result.has_value()); + ASSERT_TRUE(result->content.contains("connected")) << "the list envelope must report link state"; + EXPECT_FALSE(result->content["connected"].get()) << "the fixture endpoint never connects"; + ASSERT_TRUE(result->content.contains("timestamp")) << "the list envelope must date its values"; +} + TEST_F(OpcuaPluginTest, ListDataReturnsItems) { auto result = plugin_.list_data("tank"); ASSERT_TRUE(result.has_value()); @@ -621,6 +666,55 @@ TEST_F(OpcuaPluginAlarmsFitnessTest, IntrospectSkipsXPlcDataForAlarmsFallbackEnt EXPECT_TRUE(tank_got_x_plc_data) << "data-bearing entities must keep x-plc-data"; } +// -- Route tests -- + +namespace { + +// Calls the plugin's GET route under components/ for ``component_id`` and +// returns what the handler sent. +StubResponse get_component_status(OpcuaPlugin & plugin, const std::string & component_id) { + StubResponse recorded; + const auto routes = plugin.get_routes(); + const auto route = std::find_if(routes.begin(), routes.end(), [](const GatewayPlugin::PluginRoute & r) { + return r.method == "GET" && r.pattern.rfind("components/", 0) == 0; + }); + if (route == routes.end()) { + ADD_FAILURE() << "the plugin registers no GET route under components/"; + return recorded; + } + // The handler reads capture group 1 only. + StubRequest request{{"", component_id}}; + PluginRequest req(&request); + PluginResponse res(&recorded); + route->handler(req, res); + return recorded; +} + +} // namespace + +// The status route describes this plugin's own OPC UA session: its endpoint, +// its mode and whether it is connected. In a gateway that loads several +// plugins, other components share the entity tree, and that session says +// nothing about them. So the route serves the plugin's own component and +// answers any other one with the 404 the data route gives an entity with no +// mapped points. The own component is asked on the same plugin first, which +// shows that this harness does record a served status. +TEST_F(OpcuaPluginTest, StatusRouteServesOnlyThePluginsOwnComponent) { + ctx_.entities["other_device"] = {SovdEntityType::COMPONENT, "other_device", "/other_area", + "/other_area/other_device"}; + + const StubResponse own = get_component_status(plugin_, "test_runtime"); + ASSERT_EQ(own.status, 200) << own.body.dump(); + EXPECT_EQ(own.body.value("component_id", ""), "test_runtime"); + ASSERT_TRUE(own.body.contains("connected")) << "the plugin's own component must report link state"; + EXPECT_FALSE(own.body["connected"].get()) << "the fixture endpoint never connects"; + + const StubResponse other = get_component_status(plugin_, "other_device"); + EXPECT_EQ(other.status, 404) << "a component the plugin does not own was answered with " << other.body.dump(); + EXPECT_EQ(other.error_code, ERR_RESOURCE_NOT_FOUND); + EXPECT_FALSE(other.body.contains("endpoint_url")) << "the plugin's session must not be reported under another id"; +} + // -- FaultProvider tests -- TEST_F(OpcuaPluginTest, ListFaultsEmpty) { @@ -692,6 +786,27 @@ TEST_F(OpcuaPluginTest, GetFaultFound) { EXPECT_EQ(result->content.value("fault_code", ""), "PLC_LOW_LEVEL"); } +// A fault code is half of a record's identity, so a detail read on entity A must +// not serve entity B's record of the same code. What keeps them apart is the +// entity-scoped list this scan runs over, not a comparison inside the scan, and +// that is what this pins: two owners of one code, and only the addressed +// entity's record comes back. +TEST_F(OpcuaPluginTest, GetFaultServesOnlyTheAddressedEntitysRecordOfASharedCode) { + ctx_.all_faults = {{"faults", + {{{"fault_code", "SHARED_CODE"}, {"source_id", "other_tank"}, {"severity", 2}}, + {{"fault_code", "SHARED_CODE"}, {"source_id", "tank"}, {"severity", 3}}}}}; + + auto result = plugin_.get_fault("tank", "SHARED_CODE"); + + ASSERT_TRUE(result.has_value()); + EXPECT_EQ(result->content.value("source_id", ""), "tank") << "served another owner's record of the same code"; + EXPECT_EQ(result->content.value("severity", 0), 3); + + auto other = plugin_.get_fault("other_tank", "SHARED_CODE"); + ASSERT_TRUE(other.has_value()); + EXPECT_EQ(other->content.value("source_id", ""), "other_tank"); +} + // -- configure() validation (issue #481) -- TEST(OpcuaPluginConfigureTest, ThrowsOnInvalidNodeMap) { @@ -1158,8 +1273,358 @@ struct ScopedExecutorSpin { ScopedExecutorSpin & operator=(const ScopedExecutorSpin &) = delete; }; -// The SOVD DELETE /faults/{code} route lands on FaultProvider::clear_fault(), -// which buffers a dispatch into pending_reports_ - the SAME vector the poll +// The plugin raises under the entity it polled, so its clear has to name the +// same owner: a record is (fault_code, source_id), and a clear carrying no +// source reaches whichever record the store resolves - on a box running a second +// plugin instance against another device, that is the other device's still-active +// fault. +TEST(OpcuaPluginFaultIdentity, ClearFaultSendsTheOwningEntityAsSourceId) { + ScopedRclcpp rclcpp_scope; + auto node = std::make_shared("opcua_clear_owner_plugin"); + auto fault_manager = std::make_shared("opcua_clear_owner_faultmgr"); + + std::mutex seen_mutex; + std::vector> cleared; // (fault_code, source_id) + auto report_srv = fault_manager->create_service( + "/fault_manager/report_fault", [](const std::shared_ptr &, + const std::shared_ptr & res) { + res->accepted = true; + }); + auto clear_srv = fault_manager->create_service( + "/fault_manager/clear_fault", + [&cleared, &seen_mutex](const std::shared_ptr & req, + const std::shared_ptr & res) { + { + std::lock_guard lock(seen_mutex); + cleared.emplace_back(req->fault_code, req->source_id); + } + res->success = true; + }); + + const std::string yaml_path = "/tmp/test_opcua_clear_owner_nodemap.yaml"; + { + std::ofstream f(yaml_path); + f << R"( +area_id: owner_plc +component_id: owner_runtime +nodes: + - node_id: "ns=2;i=1" + entity_id: tank + data_name: level + data_type: float +)"; + } + + OpcuaPlugin plugin; + nlohmann::json config; + config["node_map_path"] = yaml_path; + config["endpoint_url"] = "opc.tcp://127.0.0.1:1"; // nothing listening, the fault sink drives the drain + config["poll_interval_ms"] = 100; + plugin.configure(config); + + RealNodePluginContext ctx(node.get()); + ctx.entities["tank"] = {SovdEntityType::APP, "tank", "/owner_plc", "/owner_plc/owner_runtime/tank"}; + plugin.set_context(ctx); + + ScopedExecutorSpin spinner({node, fault_manager}); + + auto probe = node->create_client("/fault_manager/report_fault"); + const auto ready_deadline = std::chrono::steady_clock::now() + std::chrono::seconds(10); + while (!probe->service_is_ready() && std::chrono::steady_clock::now() < ready_deadline) { + std::this_thread::sleep_for(std::chrono::milliseconds(20)); + } + ASSERT_TRUE(probe->service_is_ready()) << "stub ReportFault server never became discoverable"; + + // The addressed entity is the COMPONENT, the resolved owner is the hosted app. + // A record is (fault_code, owner), so the owner is what has to travel: sending + // the component id addresses a record no source owns and the fault manager + // declines it. + // Through the FaultProvider interface, which is how the gateway reaches it: a + // plugin that did not override clear_fault_record would land on the default, + // which clears by code with the addressed entity and drops the owner. + ros2_medkit_gateway::FaultProvider & provider = plugin; + static_cast(provider.clear_fault_record("owner_runtime", "SHARED_CODE", "tank")); + + const auto deadline = std::chrono::steady_clock::now() + std::chrono::seconds(10); + while (std::chrono::steady_clock::now() < deadline) { + { + std::lock_guard lock(seen_mutex); + if (!cleared.empty()) { + break; + } + } + static_cast(provider.clear_fault_record("owner_runtime", "SHARED_CODE", "tank")); + std::this_thread::sleep_for(std::chrono::milliseconds(50)); + } + + spinner.stop(); + plugin.shutdown(); + std::remove(yaml_path.c_str()); + + std::lock_guard lock(seen_mutex); + ASSERT_FALSE(cleared.empty()) << "no ClearFault reached the stub fault manager"; + EXPECT_EQ(cleared.front().first, "SHARED_CODE"); + EXPECT_EQ(cleared.front().second, "tank") + << "the clear must name the record's owner the gateway resolved, not the addressed entity"; +} + +// The gateway wires every plugin's log sink through PluginManager, which this +// suite does not link. This subclass wires one the test owns instead. +class SinkedOpcuaPlugin : public OpcuaPlugin { + public: + using GatewayPlugin::set_logger; +}; + +// Records the warnings a plugin log sink receives. The state is shared with the +// sink it hands out, so a sink copied into a callback stays valid on its own. +class WarningLog { + public: + std::function sink() const { + return [state = state_](PluginLogLevel level, const std::string & msg) { + if (level == PluginLogLevel::kWarn) { + std::lock_guard lock(state->mutex); + state->warnings.push_back(msg); + } + }; + } + + /// The first warning containing `needle`, or "" when none arrives in time. + std::string wait_for(const std::string & needle, std::chrono::milliseconds timeout) const { + const auto deadline = std::chrono::steady_clock::now() + timeout; + do { + if (auto found = find(needle); !found.empty()) { + return found; + } + std::this_thread::sleep_for(std::chrono::milliseconds(20)); + } while (std::chrono::steady_clock::now() < deadline); + return find(needle); + } + + std::string find(const std::string & needle) const { + std::lock_guard lock(state_->mutex); + for (const auto & warning : state_->warnings) { + if (warning.find(needle) != std::string::npos) { + return warning; + } + } + return ""; + } + + private: + struct State { + std::mutex mutex; + std::vector warnings; + }; + std::shared_ptr state_ = std::make_shared(); +}; + +// A clear the fault manager declined used to leave no trace: the dispatch never +// read its reply, so the REST route answered as though it had worked while the +// record stayed CONFIRMED. The reply is read now, and a refusal is logged as a +// warning that names the code and the owner the clear was addressed to. +TEST(OpcuaPluginFaultIdentity, ARefusedClearIsReadAndNotSilent) { + ScopedRclcpp rclcpp_scope; + auto node = std::make_shared("opcua_refused_clear_plugin"); + auto fault_manager = std::make_shared("opcua_refused_clear_faultmgr"); + + auto report_srv = fault_manager->create_service( + "/fault_manager/report_fault", [](const std::shared_ptr &, + const std::shared_ptr & res) { + res->accepted = true; + }); + auto clear_srv = fault_manager->create_service( + "/fault_manager/clear_fault", [](const std::shared_ptr & req, + const std::shared_ptr & res) { + res->success = false; + res->message = "Fault not found: " + req->fault_code + " (source " + req->source_id + ")"; + }); + + const std::string yaml_path = "/tmp/test_opcua_refused_clear_nodemap.yaml"; + { + std::ofstream f(yaml_path); + f << R"( +area_id: refused_plc +component_id: refused_runtime +nodes: + - node_id: "ns=2;i=1" + entity_id: tank + data_name: level + data_type: float +)"; + } + + WarningLog log; + SinkedOpcuaPlugin plugin; + plugin.set_logger(log.sink()); // before any plugin thread starts + nlohmann::json config; + config["node_map_path"] = yaml_path; + config["endpoint_url"] = "opc.tcp://127.0.0.1:1"; + config["poll_interval_ms"] = 100; + plugin.configure(config); + + RealNodePluginContext ctx(node.get()); + ctx.entities["tank"] = {SovdEntityType::APP, "tank", "/refused_plc", "/refused_plc/refused_runtime/tank"}; + plugin.set_context(ctx); + + ScopedExecutorSpin spinner({node, fault_manager}); + + auto probe = node->create_client("/fault_manager/report_fault"); + const auto ready_deadline = std::chrono::steady_clock::now() + std::chrono::seconds(10); + while (!probe->service_is_ready() && std::chrono::steady_clock::now() < ready_deadline) { + std::this_thread::sleep_for(std::chrono::milliseconds(20)); + } + ASSERT_TRUE(probe->service_is_ready()) << "stub ReportFault server never became discoverable"; + + // Addressed to the component, owned by the app: the warning has to name the + // owner the clear carried, which is not the entity in the request. + std::string warning; + const auto deadline = std::chrono::steady_clock::now() + std::chrono::seconds(10); + while (warning.empty() && std::chrono::steady_clock::now() < deadline) { + static_cast(plugin.clear_fault_record("refused_runtime", "SHARED_CODE", "tank")); + warning = log.wait_for("ClearFault refused", std::chrono::milliseconds(200)); + } + + spinner.stop(); + plugin.shutdown(); + std::remove(yaml_path.c_str()); + + ASSERT_FALSE(warning.empty()) << "a refused clear left no warning in the plugin's log"; + EXPECT_NE(warning.find("'SHARED_CODE'"), std::string::npos) << warning; + EXPECT_NE(warning.find("source 'tank'"), std::string::npos) << "the warning must name the owner: " << warning; +} + +// The ClearFault reply is handled on an executor thread whenever it arrives, +// which can be while the plugin is being destroyed, so the reply callback must +// hold nothing of the plugin. The observable half of that: the callback logs +// through the sink the plugin had when the clear was sent, and never reads the +// plugin again when the reply comes back. The stub answers only when told, the +// plugin's sink is swapped in between, and the first refusal still lands on the +// first sink. A second clear sent after the swap is the positive control: the +// second sink does receive a refusal it was wired for, so its silence about the +// first one is not a sink that hears nothing. +TEST(OpcuaPluginFaultIdentity, AClearReplyNeverReadsThePluginAgain) { + ScopedRclcpp rclcpp_scope; + auto node = std::make_shared("opcua_reply_lifetime_plugin"); + auto fault_manager = std::make_shared("opcua_reply_lifetime_faultmgr"); + + using ClearFault = ros2_medkit_msgs::srv::ClearFault; + struct Pending { + std::shared_ptr header; + std::string fault_code; + }; + std::mutex pending_mutex; + std::vector pending; + + auto report_srv = fault_manager->create_service( + "/fault_manager/report_fault", [](const std::shared_ptr &, + const std::shared_ptr & res) { + res->accepted = true; + }); + // Deferred: the request is held and answered by the test, so the reply + // arrives only after the test has changed the plugin's sink. + auto clear_srv = fault_manager->create_service( + "/fault_manager/clear_fault", + [&pending, &pending_mutex](const std::shared_ptr> & /*service*/, + const std::shared_ptr & header, + const std::shared_ptr & req) { + std::lock_guard lock(pending_mutex); + pending.push_back(Pending{header, req->fault_code}); + }); + auto answer = [&](const std::string & fault_code) { + std::lock_guard lock(pending_mutex); + for (auto it = pending.begin(); it != pending.end(); ++it) { + if (it->fault_code == fault_code) { + ClearFault::Response response; + response.success = false; + response.message = "Fault not found: " + fault_code; + clear_srv->send_response(*it->header, response); + pending.erase(it); + return true; + } + } + return false; + }; + auto wait_received = [&](const std::string & fault_code) { + const auto deadline = std::chrono::steady_clock::now() + std::chrono::seconds(10); + while (std::chrono::steady_clock::now() < deadline) { + { + std::lock_guard lock(pending_mutex); + for (const auto & p : pending) { + if (p.fault_code == fault_code) { + return true; + } + } + } + std::this_thread::sleep_for(std::chrono::milliseconds(20)); + } + return false; + }; + + const std::string yaml_path = "/tmp/test_opcua_reply_lifetime_nodemap.yaml"; + { + std::ofstream f(yaml_path); + f << R"( +area_id: lifetime_plc +component_id: lifetime_runtime +nodes: + - node_id: "ns=2;i=1" + entity_id: tank + data_name: level + data_type: float +)"; + } + + WarningLog at_dispatch; + WarningLog after_swap; + SinkedOpcuaPlugin plugin; + plugin.set_logger(at_dispatch.sink()); // before any plugin thread starts + nlohmann::json config; + config["node_map_path"] = yaml_path; + config["endpoint_url"] = "opc.tcp://127.0.0.1:1"; + config["poll_interval_ms"] = 100; + plugin.configure(config); + + RealNodePluginContext ctx(node.get()); + ctx.entities["tank"] = {SovdEntityType::APP, "tank", "/lifetime_plc", "/lifetime_plc/lifetime_runtime/tank"}; + plugin.set_context(ctx); + + ScopedExecutorSpin spinner({node, fault_manager}); + + auto probe = node->create_client("/fault_manager/report_fault"); + const auto ready_deadline = std::chrono::steady_clock::now() + std::chrono::seconds(10); + while (!probe->service_is_ready() && std::chrono::steady_clock::now() < ready_deadline) { + std::this_thread::sleep_for(std::chrono::milliseconds(20)); + } + ASSERT_TRUE(probe->service_is_ready()) << "stub ReportFault server never became discoverable"; + + static_cast(plugin.clear_fault_record("tank", "FIRST_CODE", "tank")); + ASSERT_TRUE(wait_received("FIRST_CODE")) << "the first clear never reached the stub"; + + // Stop the plugin's own threads, then swap the sink while no plugin thread + // logs. What is in flight now is only the reply to the first clear. + plugin.shutdown(); + plugin.set_logger(after_swap.sink()); + + static_cast(plugin.clear_fault_record("tank", "SECOND_CODE", "tank")); + ASSERT_TRUE(wait_received("SECOND_CODE")) << "the second clear never reached the stub"; + + ASSERT_TRUE(answer("FIRST_CODE")); + ASSERT_TRUE(answer("SECOND_CODE")); + const std::string first = at_dispatch.wait_for("'FIRST_CODE'", std::chrono::seconds(10)); + const std::string second = after_swap.wait_for("'SECOND_CODE'", std::chrono::seconds(10)); + + spinner.stop(); + std::remove(yaml_path.c_str()); + + EXPECT_FALSE(first.empty()) << "the first refusal did not reach the sink the plugin had when it sent the clear"; + EXPECT_TRUE(after_swap.find("'FIRST_CODE'").empty()) + << "the first reply read the plugin's sink when it arrived, so the callback still holds the plugin"; + EXPECT_FALSE(second.empty()) << "positive control: the swapped-in sink never received its own refusal"; + EXPECT_TRUE(at_dispatch.find("'SECOND_CODE'").empty()); +} + +// The SOVD DELETE /faults/{code} route lands on FaultProvider::clear_fault_record(), +// which buffers a dispatch into pending_reports_, the SAME vector the poll // thread drains in publish_values()/flush_pending_reports(). Before the fix the // buffer had no lock, so a push_back that reallocated the vector while another // thread iterated/swapped it corrupted the heap (the crash observed in @@ -1169,10 +1634,10 @@ struct ScopedExecutorSpin { // in send_or_buffer() and the batch.swap(pending_reports_) drain in // flush_pending_reports(). The drain only runs once report->service_is_ready() // is true, so the test stands up a real ReportFault/ClearFault server (a stub -// fault_manager) to open that gate; only then does every clear_fault() both push +// fault_manager) to open that gate. Only then does every clear both push // into the buffer AND swap it out under concurrent pushes from the other worker // threads. It must complete without heap corruption and is clean under -// ThreadSanitizer - the DDS/rclcpp machinery it drives is covered by +// ThreadSanitizer. The DDS/rclcpp machinery it drives is covered by // tsan_suppressions.txt, the same paths the gateway service tests exercise. TEST(OpcuaPluginConcurrency, ClearFaultBufferIsThreadSafe) { ScopedRclcpp rclcpp_scope; @@ -1253,7 +1718,8 @@ component_id: race_runtime // push_back against swap on the shared buffer. // The result is deliberately dropped: this test races the shared // pending-report buffer, it does not assert on any single clear. - static_cast(plugin.clear_fault("tank", "RACE_" + std::to_string(t) + "_" + std::to_string(i++ & 0x3f))); + static_cast( + plugin.clear_fault_record("tank", "RACE_" + std::to_string(t) + "_" + std::to_string(i++ & 0x3f), "tank")); } }); }