Skip to content

feat: preserve raw evidence in scan output and diagnostics - #1870

Open
mldangelo-oai wants to merge 28 commits into
mainfrom
mdangelo/codex/modelaudit-03-raw-evidence
Open

mldangelo-oai wants to merge 28 commits into
mainfrom
mdangelo/codex/modelaudit-03-raw-evidence

Conversation

@mldangelo-oai

@mldangelo-oai mldangelo-oai commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Scan reports and diagnostics now retain raw source identifiers and evidence, including embedded credentials. Remove output masking from CLI, SARIF/SBOM, CatBoost and selected detector previews while retaining evidence bounds, detections and failure classification.

This is PR 3 of the six-PR simplification stack, now targeting main after PRs 1 and 2 landed. Normalization remains only where it determines finding identity, grouping or classification. Stream, directory-owner, Hugging Face and MLflow producers retain those inputs separately from raw fields, including across aggregation and saved JSON. Direct SARIF and SBOM APIs retain their original classification behavior. Producer metadata carries the original SBOM type/risk inputs through emitted raw source keys and saved JSON; the CLI retains its historical source classification.

file_metadata now uses emitted source keys: ordinary sources stay raw, and oversized source identifiers use bounded preview/digest identifiers shared by keys and references. Consumers must use those keys; previously masked keys are no longer lookup aliases. Metadata values and attribution among distinct sources are preserved. README and changelog document the intentional output/API change. The exported redact_huggingface_url_for_display and redact_huggingface_urls_in_text helpers are retired; callers that need masking must apply their own presentation policy. Debug help and output tell users to inspect raw diagnostics before sharing.

Validation:

Current-main integration (2a5185a): preserves the upstream SafeTensors FDICT fix and deterministic initial/retried cache interruption fixture. Focused cache/routing checks and five neighboring compression cases pass on this branch; real benign/malicious CLI scans preserve scanner selection, JSON findings, SHA-256 metadata, and exit codes 0/1. Ruff checks pass. Changed-file mypy passes; full local mypy retains the previously documented optional TensorFlow baseline error. Hosted CI validates the published head.

  • Both interruption cleanup tests control initial/retried capture deterministically and verify that every acquired monitor closes. Retained-interruption cases pass on each current prefix; leak mutations are rejected.

  • The latest base integration retains the reviewed scanner/test changes and current released dependency floor. A fresh-process logging test fixes deterministic CLI-module pollution exposed by Windows CI; all 33 ordered logging/debug/interruption cases pass on each updated prefix. The released dependency and evidence/debug cohort also passes all 368 cases per prefix. The upstream cache retry regression is included unchanged.

  • Final review follow-up moves report implementation details into contributor documentation and annotates the renamed JFrog regression. The earlier AST audit confirmed all 247 then-changed/new test functions were typed; 13 JFrog CLI regressions pass on each corrected prefix. Seven paired actual CLI cases preserve parent behavior for malformed URLs and bracketed local paths.

  • Follow-up validation: 347 focused tests pass on each corrected PR3–6 prefix. All 22,852 collected root-suite node IDs fit the Windows environment-variable limit after assigning short IDs to the two large-input tests; their inputs/assertions are unchanged. Real debug CLI, JSON and help invocations pass. Twenty-four baseline/candidate Hugging Face exception-classification and inventory cases match.

  • Before the portability/documentation follow-up, the committed prefix passed 1,948 affected tests, with three Windows-only and two opt-in integration skips. Import and native-artifact checks confirmed this checkout. Full Ruff format/lint passed; mypy checked 495 files and reports only the existing unreachable optional TensorFlow import at tests/scanners/test_weight_distribution_scanner.py:2122.

  • Parent/candidate comparisons cover 150 MLflow failure cases, 283 stream identities, and eight real JAX/TensorFlow directory-owner variants. Independent public-path probes verify malicious pickle stream aggregation, actual MLflow/Hugging Face CLI failures, saved-result identities, rule tags, grouping/counts and SBOM risk attribution.

  • 333 SBOM comparisons preserve the distinction between literal public API inputs and CLI source normalization. Independent checks cover 28 direct type/size/hash cases and 36 CLI path-tracking cases, plus distinct signed sources with different content hashes.

  • Seventeen additional end-to-end cases pass unchanged on the parent and candidate, covering public emitted-asset/metadata exports, stream failures, empty/incomplete scans, Hugging Face errors and actual CLI auto-streaming. Historical risk grouping of query variants remains intact.

  • Eighteen final regressions also pass unchanged on the parent: long MLflow sources, bounded display collisions, literal versus emitted-path SBOM exports, and malformed-Unicode SARIF paths. Classification precedes display truncation.

  • Bounded MLflow locations remain distinct through saved JSON, including repeated/reordered inputs and per-source risk attribution. Stream, cloud, JFrog and retry terminal outputs filter control characters while preserving raw result/error evidence. New parent comparisons cover the saved-result and terminal contracts, including cache rejection warnings and JFrog debug output. Hugging Face probe diagnostics also retain terminal formatting.

  • Thirty-three parent-derived regression cases cover oversized-source serialization joins, literal alias collisions, and Hugging Face failures. Source-owned exceptions preserve their original classification, type, cause, and string-conversion count while reporting raw evidence. Direct and saved-result SBOM exports agree.

  • Terminal diagnostics escape CR/LF/tab as well as other unsafe controls; evidence previews retain their existing multiline contract. Parent/candidate probes cover quoted MLflow credentials and diagnostic output without changing acquired sources.

  • Oversized finding text retains its historical bound and SARIF fingerprint through JSON reload; an actual TFLite custom-operator scan verifies the original fingerprint. Shared source IDs preserve artifact attribution through direct SARIF and saved JSON, including literal identifiers that collide after path or URI normalization. Bounded previews escape unpaired Unicode characters while hashing the original source. JSON export remains available when the working directory has been removed or is inaccessible; Artificial source identifiers are explicitly relative, and conservative basename reservation keeps saved sources distinct across directory changes; actual SARIF rendering retains its original errors. Nested evidence retains its original depth budget. Hugging Face dry-run metadata failures keep the original failure category.

  • Earlier paired CLI runs cover clean, malicious, corrupt and missing inputs with exit codes 0/1/2. The final correction adds real producer/CLI regressions; no assertions were weakened to accept changed detection or fail-closed behavior.

Local full-suite, cross-platform and live remote acquisition completion is not claimed. One live Hugging Face inventory test requires unavailable SOCKS support (socksio); the same failure reproduces on the unchanged parent. External transport is mocked in deterministic remote-source tests. Upstream-pinned mypy 2.4 could not be installed through this devbox's index; local checks use 2.3.1 with the installed Python 3.12 profile.

Measured reduction: 9,622 maintained lines — production −4,352; tests −5,306; documentation +36. Counts include all retained identity helpers and added regression tests. Generated code is unchanged.

PRs 1: tooling and 2: tests have landed. Remaining review order: 3: raw evidence → 4: picklescan → 5: scanning/results → 6: acquisition/cache/progress. PR3 targets main; PRs 4–6 target the preceding branch.

The canonical agent guide explicitly preserves raw local evidence and prohibits reintroducing credential redaction, while retaining terminal escaping, permissions, and detections.

The README explicitly documents that the legacy redacted_value field remains for compatibility and contains bounded raw evidence without masking guarantees. Detector and passive-sidecar schemas remain consistent; all 96 focused legacy-field cases pass.

The security policy, canonical agent guide, user security model, and CVE triage guide explicitly define unredacted local output as supported behavior and accepted risk. Missing masking is not a vulnerability by itself and must not be repaired by reintroducing redaction. This includes secrets and credential-bearing URLs in local results, reports, logs, and cache metadata. Unauthorized host-data access or transmission remains in scope; telemetry collection limits are unchanged. This policy follow-up changes documentation only and passes formatting/link checks.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-03T22:22:17.330291Z d771317 New commits
🔒 Security Review ✅ Completed 2026-10-03T22:21:15.824553Z d771317 New commits

Security findings

Advisory findings (3)

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Workflow run and artifacts

Performance Benchmarks

Compared 13 shared benchmarks with a regression threshold of 15%.
Status: 0 regressions, 0 improved, 13 stable, 0 new, 0 missing.
Aggregate shared-benchmark median: 3.909s -> 3.881s (-0.7%).

Workload Benchmark Target Size Files Baseline Current Change Status
warm-cache-rescan tests/benchmarks/test_scan_benchmarks.py::test_scan_warm_cached_repository_rescan release-candidate 547.3 KiB 32 125.79ms 130.13ms +3.5% stable
mixed-model-repository tests/benchmarks/test_scan_benchmarks.py::test_scan_release_candidate_repository release-candidate 547.3 KiB 32 610.75ms 594.95ms -2.6% stable
duplicate-heavy-registry tests/benchmarks/test_scan_benchmarks.py::test_scan_duplicate_registry_snapshot registry-snapshot 915.2 KiB 13 564.31ms 550.04ms -2.5% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_raw] nested_raw 78 B 1 253.2us 248.3us -1.9% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_hex] nested_hex 130 B 1 281.8us 284.5us +0.9% stable
padded-multi-stream-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_padded_multi_stream_upload multi_stream_padded 4.1 KiB 1 298.3us 299.5us +0.4% stable
clean-training-checkpoint tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_clean_training_checkpoint safe_large 278.2 KiB 1 110.08ms 110.34ms +0.2% stable
suspicious-pickle-intake tests/benchmarks/test_scan_benchmarks.py::test_scan_suspicious_pickle_intake suspicious-intake 183.8 KiB 4 120.43ms 120.19ms -0.2% stable
single-checkpoint-preflight tests/benchmarks/test_scan_benchmarks.py::test_scan_single_checkpoint_before_load single_checkpoint.pkl 183.0 KiB 1 95.97ms 96.08ms +0.1% stable
rejected-basic-auth-candidates tests/benchmarks/test_scan_benchmarks.py::test_rejected_basic_auth_candidates_scan_linearly - 371.1 KiB 1 2.167s 2.165s -0.1% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_base64] nested_base64 98 B 1 275.6us 275.8us +0.1% stable
direct-malicious-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_direct_malicious_upload malicious_reduce 52 B 1 202.1us 202.0us -0.1% stable
chunked-upload-stream tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_chunked_upload_stream chunked_stream 278.2 KiB 1 113.32ms 113.33ms +0.0% stable

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b863549efd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread modelaudit/cli.py
audit_result,
_classification_paths={
path: _cli_source_classification_path(
path if is_mlflow_uri(path) else _source_identity_path(path, audit_result.file_metadata.get(path))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Map long MLflow sources to their emitted SBOM identity

When --sbom is requested after an MLflow acquisition refusal and the model URI exceeds 512 characters, scan_mlflow_model() stores the issue and file_metadata under _mlflow_report_source_identifier(model_uri) (the bounded digest-bearing key), while paths_for_sbom still contains the original URI. Passing that raw URI here means generate_sbom_pydantic() cannot find its metadata and compares the issue against a different classification path, so the resulting component loses its acquisition metadata and reports risk_score=0 instead of the informational finding's score; _report_source_path() also produces a different ellipsis-only reference. Resolve the raw URI to the emitted MLflow key/source identity before generating the component.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The CLI still supplies separate classification and presentation overrides in the SBOM writer. Twelve parent/current refusal comparisons across all three builders preserve CLI risk 1 and identical component properties/types; raw direct-API risk 0 occurs in both versions, while emitted finding paths retain risk 1. The committed test_cli_mlflow_sbom_classifies_source_before_display_bound and literal-input/emitted-path API regressions exercise these distinctions; all 13 MLflow SBOM cases passed. The original refusal builders had no acquisition file metadata to preserve. Mapping raw API inputs to emitted identities would change this baseline contract, so no additional mapping is needed.

@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Hooray!

Reviewed commit: b863549efd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 👍

Reviewed commit: 4b40b19ddd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 4b40b19ddd

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Hugging Face classification re-renders wrapped exceptions, breaking the added single-conversion regression.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)

Comment thread modelaudit/utils/sources/huggingface.py
@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: cb7e029a8b

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Raw Hugging Face cause text can still alter access classification through credential-like status markers.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cb7e029a8b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines 89 to 92
logger.debug(
f"Attempt {attempt + 1} failed for {getattr(func, '__name__', 'unknown')}: {display_error}. "
f"Attempt {attempt + 1} failed for {getattr(func, '__name__', 'unknown')}: "
f"{format_terminal_text(str(e))}. "
f"Retrying in {delay:.1f} seconds..."

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Honor the retry error sanitizer

When a caller supplies sanitize_error, this debug log still renders the raw exception because neither this wrapper nor the decorator factory invokes the callback. In verbose/debug environments, a credential-bearing provider exception or signed URL can therefore be written to logs despite the caller explicitly supplying a masking policy; the still-accepted parameter has silently become a no-op. Apply the callback before terminal formatting and continue forwarding it through retry_with_backoff and retry_cloud_operation.

AGENTS.md reference: AGENTS.md:L23-L26

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The callback's masking semantics are intentionally retired under the raw local evidence policy. The compatibility comment records why the keyword remains accepted. The retained tests check all three entry-point signatures, retry attempts, original exception identity, and raw single-line diagnostics. A failure-path probe also passed a masking callback to each entry point: each retried twice, retained the original error, made zero callback calls, and escaped the synthetic newline. This is keyword compatibility, not preservation of the former masking callback behavior.

Comment on lines 988 to 991
return {
"url_sha256": self.get_cache_key(url),
"url_display": redact_url_for_display(url),
"url_display": url,
"url_scheme": parsed.scheme,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep signed URLs out of cache metadata

When the default cloud cache is used with a presigned or userinfo-bearing URL, this value is persisted verbatim in cache_metadata.json by cache_file(), leaving reusable credentials on disk after the scan. The SHA-256 field already provides the query-sensitive cache identity and url_display is not read anywhere, so retain a credential-redacted display value rather than the raw URL; restrictive file permissions do not protect copies in backups or diagnostic bundles.

AGENTS.md reference: AGENTS.md:L23-L26

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Raw url_display is intentional local cache metadata under the raw local evidence policy; the policy is explicit in commit 77b8963ad195f3ba228ad2566e63e2abd2a9e60e. URL metadata retains query-sensitive SHA-256 identity, while metadata permissions remain private on POSIX. Existing cache tests verify signed-URL retention, identity, permissions, and migration of legacy entries. The credential-retention observation is correct; restoring masking would reverse the agreed behavior.

@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 77b8963ad1

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 🚀

Reviewed commit: 77b8963ad1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Auth classification can inspect raw credential text, and saved oversized source identities can lose SBOM component type.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Preserve source type when serializing oversized identifiers

modelaudit/​integrations/​sbom_generator.py:149

An oversized stream or Hugging Face classification path does not survive a saved-result round trip. serialize_source_value() replaces a source identity over 256 KiB with a modelaudit-source:...<sha256> identifier; returning that identifier here means _get_component_type() can no longer see the original suffix. For example, a long URL ending in model.pkl is classified as a model in memory but as a generic file after JSON reload. Persist a bounded classification suffix/type separately from the emitted identifier and use it here.

@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

Please review the current head after integration with main 2a5185a. The new changes preserve the upstream SafeTensors FDICT fix and deterministic cache cleanup regressions. Existing review-thread replies document validated fixes, parent comparisons, and the explicit raw local evidence policy in AGENTS.md. Please verify current source and avoid repeating already addressed or intentional behavior.

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 1435481b59

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Raw credentials are assigned to the misleading public field redacted_value, creating an unsafe API contract for consumers.

Review effort: Balanced
Findings: 2 Medium severity

Open (2)
Resolved since last review (3)

"confidence": round(confidence, 2),
"pattern": pattern.pattern[:50] + "..." if len(pattern.pattern) > 50 else pattern.pattern,
"redacted_value": "Basic <redacted>",
"redacted_value": format_evidence_string(matched_text),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Documented the legacy field explicitly in the public report contract: redacted_value retains its name for compatibility, but now contains bounded raw evidence and provides no masking guarantee. The serialized key and detection behavior are preserved. The focused detector/text/passive-sidecar cohort passes all 96 cases, including raw Basic credentials and bounded previews.

"confidence": 0.8,
"pattern": "passive_data_auth_line",
"redacted_value": "Basic <redacted>",
"redacted_value": format_evidence_string("Basic " + token[:180].decode("ascii")),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The README now explicitly documents the same legacy redacted_value contract for detector and passive-sidecar findings: retained key, bounded raw evidence, no masking guarantee. Both producers keep that schema. The committed passive Basic-auth finding and bounded-preview tests pass as part of the 96-case focused cohort; no runtime rename or alias is needed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants