Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions monitoring/uss_qualifier/configurations/configuration.py
Original file line number Diff line number Diff line change
Expand Up @@ -255,6 +255,10 @@ class TimingReportConfiguration(ImplicitDict):
"""Percentage of test time to break down in the timing report (smaller contributions are not reported)"""


class AIHelpersConfiguration(ImplicitDict):
pass


class ArtifactsConfiguration(ImplicitDict):
raw_report: Optional[RawReportConfiguration] = None
"""Configuration for raw report generation"""
Expand All @@ -277,6 +281,9 @@ class ArtifactsConfiguration(ImplicitDict):
timing_report: Optional[TimingReportConfiguration] = None
"""If specified, configuration describing a desired report describing where and how time was spent during the test."""

ai_helpers: Optional[AIHelpersConfiguration] = None
"""If specified, configuration describing AI helper tools to include in artifacts."""

@property
def acceptable_findings(self) -> Iterable[FullyQualifiedCheck]:
"""Iterates through checks where findings are acceptable in at least one tested_requirements artifact."""
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -438,6 +438,9 @@ v1:
timing_report:
percentage_of_time_to_break_down: 90

# Write out helper resources for AI agents interpreting the results
ai_helpers: {}

# This block defines whether to return an error code from the execution of uss_qualifier, based on the content of the
# test run report. All of the criteria must be met to return a successful code.
validation:
Expand Down
29 changes: 29 additions & 0 deletions monitoring/uss_qualifier/reports/AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
This folder contains definitions for artifacts generated from the output of a uss_qualifier test report.

If your user has requested interpretation, troubleshooting, diagnosis, or similar for a uss_qualifier test report, the instructions below are intended to help with that task.

### Background and objective

uss_qualifier is an automated testing tool developed by InterUSS in this repository (`monitoring`) for the purpose of qualifying systems designed to conduct UTM (UAS traffic management) activities. A uss_qualifier test run produces a set of artifacts generally containing a subset of the report types defined here. You will be using these artifacts to determine why a particular outcome occurred, such as a participant being marked as failing or "Not fully verified" for a set of tested requirements. Your investigation should seek to identify the root causes as deeply as possible with the available information. A detailed analysis of every failed check is generally undesirable; instead, you should seek to identify the patterns behind multiple check failures and the likely cause(s) of those patterns. Ideally, your analysis should conclude with one or more top-level suggestions for what participant(s) should change to obtain a fully-passing test result in the future.

### Investigation procedure

Your user should have provided you with a zip file of artifacts, or pointed you to a folder containing artifacts, or provided report.json; if this is not the case, indicate to the user that you require these artifacts to perform the requested task according to these instructions and stop your task unless requested otherwise. The primary artifact is report.json which is a `TestRunReport` as defined in report.py of this folder. Note that the version of uss_qualifier that conducted the test run may be different than the codebase version containing these instructions. It should, however, still indicate which codebase version was used in report.codebase_version. If your user has provided a `monitoring` folder with a clone of the InterUSS `monitoring` repository in your workspace, prepare for this investigation by checking out the specified version of uss_qualifier with that repository or asking your user to do this (but take care not to lose any work your user may have been performing in that repository). If you have a local clone of the `monitoring` repository checked out to the appropriate version, find files and other resources linked from this document in that clone rather than relative to the HTTP location you may have retrieved this document from.

In most investigations, the tested requirements artifact is generally the first artifact to inspect. status.json provides a machine-readable top-level summary of the status of each participant with regard to the requirements intended to be verified by the test configuration. Failed or not-fully-verified statuses are usually observations to investigate further. To troubleshoot these observations, refer to the instructions in [the tested_requirements troubleshooting documentation](./tested_requirements/README.md). When following these instructions, remember that test scenario documentation may be found in the `monitoring` folder in your workspace if available rather than needing to make HTTP requests. Following those tested_requirements troubleshooting instructions (including moving to inspection of portions of the sequence view artifact) should allow you to form a holistic understanding of the intent and expectations of each test and what differed in the test run under investigation versus nominal outcomes.

### Responses

In your response, refer users to key specific observations that lead to your conclusions (or example specific observations that illustrate your conclusions) when possible. When these observations are events that happened during the test run, reference the description of these events in the sequence view artifact. The shorthand name for a specific event is "sNeM" (for instance, s14e257) referring to the Nth scenario's Mth event, and the sequence view scenario pages contain linkable anchors to each event.

### Special considerations

#### DSS behavior

If troubleshooting involves the behavior of DSS (discovery and synchronization service) instances and your user has provided a `dss` folder with a clone of the InterUSS `dss` repository, this may also be a valuable resource for understanding observed behavior. Certain portions of the test artifacts may provide information about which version(s) of the DSS were deployed by DSS instance providers, or you may consider asking your user if they know which DSS versions were deployed if this is likely to be important information that you cannot otherwise determine. DSS behavior can be inspected in the `dss` repository after it is checked out to the appropriate version.

#### Execution errors

A common problem in failed uss_qualifier test runs is an execution error in uss_qualifier. As noted in [the sequence view documentation](./sequence_view/README.md), execution errors always indicate a uss_qualifier bug and this should be surfaced to your user. If possible, you should attempt to determine whether this bug has been previously reported and/or fixed by checking [GitHub Issues for the `monitoring` repo](https://github.com/interuss/monitoring/issues). If the bug has already been reported, your response should mention and/or link to the existing Issue. If the bug has already been fixed in an existing release, you should note that this problem can be resolved by updating to the more recent codebase version. If this bug has not yet been reported, please consider offering to draft a bug report using [the appropriate template](https://github.com/interuss/monitoring/issues/new?template=bug_test_scenario.md) unless the user has specifically requested brevity or similar. InterUSS finds these bug reports very useful and submitting them for previously-unidentified bugs helps the entire community.

In addition to the above, however, also note to yourself that the majority of execution errors are surfaced because an unusual/untested uss_qualifier code path has been taken because one or more participants have failed in ways not yet exercised by InterUSS. For instance, many check-failure error message strings are never generated in InterUSS CI tests because there are so many different ways for systems under test to fail. So, if you are investigating a test run that includes an execution error, an important part of your investigation should be to determine why that bug was elicited in this test run and, presumably, not in more nominal test runs. The answer to this question is frequently that one or more participants have one or more underlying failures. In this case, those underlying failures should be emphasized as the most interesting root causes since fixing the InterUSS bug would merely reveal these underlying failures more clearly and cleanly.
11 changes: 10 additions & 1 deletion monitoring/uss_qualifier/reports/README.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,15 @@
# uss_qualifier reports

## Troubleshooting test runs

The recommended approach for understanding and troubleshooting a particular test run of uss_qualifier with a lower level of subject-matter expertise is to use an AI coding agent. Ideally, prepare for the prompt by cloning this repository (and possibly [the `dss` repository](https://github.com/interuss/dss)) into the workspace accessible by the coding agent; this allows the coding agent to easily examine specific behaviors of uss_qualifier (and perhaps the DSS) in its investigation. Place the artifacts produced by the test run (folder, .zip, or report.json) into the workspace accessible by the coding agent as well. The suggested form for the prompt is to briefly describe your objective and refer the coding agent to the latest version of agent test result interpretation instructions ([AGENTS.md](./AGENTS.md)) even if the version of uss_qualifier used in the test run was older. Examples:

> Please use the instructions at https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/reports/AGENTS.md to help me understand why the tested requirements in the test run artifacts located in test_runs/683d882f-17b6-4ca5-8726-81612fa96cad indicate that Example USS failed.

> Please use the instructions at https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/reports/AGENTS.md to help me understand why no one seems to be passing in the artifacts in test_runs/683d882f-17b6-4ca5-8726-81612fa96cad.

> Using the instructions at https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/reports/AGENTS.md, why are there so many failures in scenario 11 in the artifacts in test_runs/683d882f-17b6-4ca5-8726-81612fa96cad?

## Report types

uss_qualifier is capable of generating a range of artifacts from a test run, each intended to fulfill a different purpose. Part of a [test configuration](../configurations) defines artifacts that should be produced by the test run.
Expand Down Expand Up @@ -29,4 +39,3 @@ The [globally-expanded report artifact](./globally_expanded/README.md) assembles
### Test artifacts obfuscation tool

The [obfuscation tool](./obfuscate.md) can be used to redact and pseudo-anonymize participant IDs, server hostnames, and authorization tokens from a collection of test artifacts before sharing or publishing.

Empty file.
34 changes: 34 additions & 0 deletions monitoring/uss_qualifier/reports/ai_helpers/ai_helpers.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
import os

from monitoring.uss_qualifier.configurations.configuration import ArtifactsConfiguration
from monitoring.uss_qualifier.reports import jinja_env


def write_ai_helper_files(output_path: str, artifacts: ArtifactsConfiguration) -> None:
# Ensure output_path exists
os.makedirs(output_path, exist_ok=True)

# Extract subfolder names for tested requirements configurations
tested_requirements_folders = (
[cfg.report_name for cfg in artifacts.tested_requirements]
if artifacts.tested_requirements
else []
)

# Render and write AGENTS.md
agents_template = jinja_env.get_template("ai_helpers/AGENTS.md.template")
agents_content = agents_template.render()
with open(os.path.join(output_path, "AGENTS.md"), "w") as f:
f.write(agents_content)

# Render and write README.md
readme_template = jinja_env.get_template("ai_helpers/README.md.template")
readme_content = readme_template.render(
has_raw_report=(artifacts.raw_report is not None),
tested_requirements_folders=tested_requirements_folders,
has_timing_report=(artifacts.timing_report is not None),
has_globally_expanded_report=(artifacts.globally_expanded_report is not None),
has_report_html=(artifacts.report_html is not None),
)
with open(os.path.join(output_path, "README.md"), "w") as f:
f.write(readme_content)
15 changes: 15 additions & 0 deletions monitoring/uss_qualifier/reports/artifacts.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
from loguru import logger

from monitoring.uss_qualifier.configurations.configuration import ArtifactsConfiguration
from monitoring.uss_qualifier.reports.ai_helpers.ai_helpers import write_ai_helper_files
from monitoring.uss_qualifier.reports.documents import make_report_html
from monitoring.uss_qualifier.reports.globally_expanded.generate import (
generate_globally_expanded_report,
Expand Down Expand Up @@ -145,6 +146,19 @@ def make_timing_report() -> None:
generate_timing_report(redacted_report, artifacts.timing_report, path)
logger.info(f"Wrote timing report in {time.monotonic() - t0:.1f}s")

def make_ai_helpers() -> None:
if artifacts.ai_helpers is None:
return
if artifacts.sequence_view is None or not artifacts.tested_requirements:
logger.warning(
"AI helper files generation requested, but either sequence_view or tested_requirements is not configured. Disabling AI helpers."
)
return
logger.info(f"Writing AI troubleshooting helper files to {output_path}")
t0 = time.monotonic()
write_ai_helper_files(output_path, artifacts)
logger.info(f"Wrote AI helper files in {time.monotonic() - t0:.1f}s")

artifact_generators = [
make_raw_report,
make_html_report,
Expand All @@ -153,6 +167,7 @@ def make_timing_report() -> None:
make_sequence_view,
make_globally_expanded_report,
make_timing_report,
make_ai_helpers,
]
generators = [Process(target=g, daemon=True) for g in artifact_generators]
for p in generators:
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
# Instructions for AI Assistants & Agents

If you are an AI assistant (like ChatGPT, Claude, Gemini, Grok) troubleshooting this test report, refer to the instructions at https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/reports/AGENTS.md. If you cannot access the content at that location, request that your user provide that file to you directly.
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# uss_qualifier Test Run Report Artifacts

This directory contains the artifacts generated from a uss_qualifier test run.

## Artifacts Layout:
{% if has_raw_report %}
- **`report.json`**: The raw TestRunReport.
{% endif %}
{% if tested_requirements_folders %}
- **Tested Requirements Views**: Summary of tested requirements per participant:
{% for folder in tested_requirements_folders %}
- `{{ folder }}/`: Contains HTML compliance reports and `status.json` for participant verifications.
{% endfor %}
{% endif %}
- **`sequence/`**: Chronological log of events and scenarios.
{% if has_timing_report %}
- **`timing/`**: Test execution timing details.
{% endif %}
{% if has_globally_expanded_report %}
- **`globally_expanded/`**: Mimics manual test run logs.
{% endif %}
{% if has_report_html %}
- **`report.html`**: Human-readable HTML layout of the raw report.
{% endif %}

---

## 🤖 For AI Assistants & Agents
If you are an AI assistant (like ChatGPT, Claude, Gemini, Grok) analyzing this report, please read **[AGENTS.md](AGENTS.md)** first. It contains instructions on how to investigate and explain the results to the user.
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
{
"$id": "https://github.com/interuss/monitoring/blob/main/schemas/monitoring/uss_qualifier/configurations/configuration/AIHelpersConfiguration.json",
"$schema": "https://json-schema.org/draft/2020-12/schema",
"description": "monitoring.uss_qualifier.configurations.configuration.AIHelpersConfiguration, as defined in monitoring/uss_qualifier/configurations/configuration.py",
"properties": {
"$ref": {
"description": "Path to content that replaces the $ref",
"type": "string"
}
},
"type": "object"
}
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,17 @@
"description": "Path to content that replaces the $ref",
"type": "string"
},
"ai_helpers": {
"description": "If specified, configuration describing AI helper tools to include in artifacts.",
"oneOf": [
{
"type": "null"
},
{
"$ref": "AIHelpersConfiguration.json"
}
]
},
"globally_expanded_report": {
"description": "If specified, configuration describing a desired report mimicking what might be seen had the test run been conducted manually.",
"oneOf": [
Expand Down
Loading