You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Scan reports and diagnostics now retain raw source identifiers and evidence, including embedded credentials. Remove output masking from CLI, SARIF/SBOM, CatBoost and selected detector previews while retaining evidence bounds, detections and failure classification.
This is PR 3 of the six-PR simplification stack, now targeting main after PRs 1 and 2 landed. Normalization remains only where it determines finding identity, grouping or classification. Stream, directory-owner, Hugging Face and MLflow producers retain those inputs separately from raw fields, including across aggregation and saved JSON. Direct SARIF and SBOM APIs retain their original classification behavior. Producer metadata carries the original SBOM type/risk inputs through emitted raw source keys and saved JSON; the CLI retains its historical source classification.
file_metadata now uses emitted source keys: ordinary sources stay raw, and oversized source identifiers use bounded preview/digest identifiers shared by keys and references. Consumers must use those keys; previously masked keys are no longer lookup aliases. Metadata values and attribution among distinct sources are preserved. README and changelog document the intentional output/API change. The exported redact_huggingface_url_for_display and redact_huggingface_urls_in_text helpers are retired; callers that need masking must apply their own presentation policy. Debug help and output tell users to inspect raw diagnostics before sharing.
Validation:
Current-main integration (2a5185a): preserves the upstream SafeTensors FDICT fix and deterministic initial/retried cache interruption fixture. Focused cache/routing checks and five neighboring compression cases pass on this branch; real benign/malicious CLI scans preserve scanner selection, JSON findings, SHA-256 metadata, and exit codes 0/1. Ruff checks pass. Changed-file mypy passes; full local mypy retains the previously documented optional TensorFlow baseline error. Hosted CI validates the published head.
Both interruption cleanup tests control initial/retried capture deterministically and verify that every acquired monitor closes. Retained-interruption cases pass on each current prefix; leak mutations are rejected.
The latest base integration retains the reviewed scanner/test changes and current released dependency floor. A fresh-process logging test fixes deterministic CLI-module pollution exposed by Windows CI; all 33 ordered logging/debug/interruption cases pass on each updated prefix. The released dependency and evidence/debug cohort also passes all 368 cases per prefix. The upstream cache retry regression is included unchanged.
Final review follow-up moves report implementation details into contributor documentation and annotates the renamed JFrog regression. The earlier AST audit confirmed all 247 then-changed/new test functions were typed; 13 JFrog CLI regressions pass on each corrected prefix. Seven paired actual CLI cases preserve parent behavior for malformed URLs and bracketed local paths.
Follow-up validation: 347 focused tests pass on each corrected PR3–6 prefix. All 22,852 collected root-suite node IDs fit the Windows environment-variable limit after assigning short IDs to the two large-input tests; their inputs/assertions are unchanged. Real debug CLI, JSON and help invocations pass. Twenty-four baseline/candidate Hugging Face exception-classification and inventory cases match.
Before the portability/documentation follow-up, the committed prefix passed 1,948 affected tests, with three Windows-only and two opt-in integration skips. Import and native-artifact checks confirmed this checkout. Full Ruff format/lint passed; mypy checked 495 files and reports only the existing unreachable optional TensorFlow import at tests/scanners/test_weight_distribution_scanner.py:2122.
Parent/candidate comparisons cover 150 MLflow failure cases, 283 stream identities, and eight real JAX/TensorFlow directory-owner variants. Independent public-path probes verify malicious pickle stream aggregation, actual MLflow/Hugging Face CLI failures, saved-result identities, rule tags, grouping/counts and SBOM risk attribution.
333 SBOM comparisons preserve the distinction between literal public API inputs and CLI source normalization. Independent checks cover 28 direct type/size/hash cases and 36 CLI path-tracking cases, plus distinct signed sources with different content hashes.
Seventeen additional end-to-end cases pass unchanged on the parent and candidate, covering public emitted-asset/metadata exports, stream failures, empty/incomplete scans, Hugging Face errors and actual CLI auto-streaming. Historical risk grouping of query variants remains intact.
Eighteen final regressions also pass unchanged on the parent: long MLflow sources, bounded display collisions, literal versus emitted-path SBOM exports, and malformed-Unicode SARIF paths. Classification precedes display truncation.
Bounded MLflow locations remain distinct through saved JSON, including repeated/reordered inputs and per-source risk attribution. Stream, cloud, JFrog and retry terminal outputs filter control characters while preserving raw result/error evidence. New parent comparisons cover the saved-result and terminal contracts, including cache rejection warnings and JFrog debug output. Hugging Face probe diagnostics also retain terminal formatting.
Thirty-three parent-derived regression cases cover oversized-source serialization joins, literal alias collisions, and Hugging Face failures. Source-owned exceptions preserve their original classification, type, cause, and string-conversion count while reporting raw evidence. Direct and saved-result SBOM exports agree.
Terminal diagnostics escape CR/LF/tab as well as other unsafe controls; evidence previews retain their existing multiline contract. Parent/candidate probes cover quoted MLflow credentials and diagnostic output without changing acquired sources.
Oversized finding text retains its historical bound and SARIF fingerprint through JSON reload; an actual TFLite custom-operator scan verifies the original fingerprint. Shared source IDs preserve artifact attribution through direct SARIF and saved JSON, including literal identifiers that collide after path or URI normalization. Bounded previews escape unpaired Unicode characters while hashing the original source. JSON export remains available when the working directory has been removed or is inaccessible; Artificial source identifiers are explicitly relative, and conservative basename reservation keeps saved sources distinct across directory changes; actual SARIF rendering retains its original errors. Nested evidence retains its original depth budget. Hugging Face dry-run metadata failures keep the original failure category.
Earlier paired CLI runs cover clean, malicious, corrupt and missing inputs with exit codes 0/1/2. The final correction adds real producer/CLI regressions; no assertions were weakened to accept changed detection or fail-closed behavior.
Local full-suite, cross-platform and live remote acquisition completion is not claimed. One live Hugging Face inventory test requires unavailable SOCKS support (socksio); the same failure reproduces on the unchanged parent. External transport is mocked in deterministic remote-source tests. Upstream-pinned mypy 2.4 could not be installed through this devbox's index; local checks use 2.3.1 with the installed Python 3.12 profile.
Measured reduction: 9,622 maintained lines — production −4,352; tests −5,306; documentation +36. Counts include all retained identity helpers and added regression tests. Generated code is unchanged.
The canonical agent guide explicitly preserves raw local evidence and prohibits reintroducing credential redaction, while retaining terminal escaping, permissions, and detections.
The README explicitly documents that the legacy redacted_value field remains for compatibility and contains bounded raw evidence without masking guarantees. Detector and passive-sidecar schemas remain consistent; all 96 focused legacy-field cases pass.
The security policy, canonical agent guide, user security model, and CVE triage guide explicitly define unredacted local output as supported behavior and accepted risk. Missing masking is not a vulnerability by itself and must not be repaired by reintroducing redaction. This includes secrets and credential-bearing URLs in local results, reports, logs, and cache metadata. Unauthorized host-data access or transmission remains in scope; telemetry collection limits are unchanged. This policy follow-up changes documentation only and passes formatting/link checks.
The reason will be displayed to describe this comment to others. Learn more.
Map long MLflow sources to their emitted SBOM identity
When --sbom is requested after an MLflow acquisition refusal and the model URI exceeds 512 characters, scan_mlflow_model() stores the issue and file_metadata under _mlflow_report_source_identifier(model_uri) (the bounded digest-bearing key), while paths_for_sbom still contains the original URI. Passing that raw URI here means generate_sbom_pydantic() cannot find its metadata and compares the issue against a different classification path, so the resulting component loses its acquisition metadata and reports risk_score=0 instead of the informational finding's score; _report_source_path() also produces a different ellipsis-only reference. Resolve the raw URI to the emitted MLflow key/source identity before generating the component.
The reason will be displayed to describe this comment to others. Learn more.
The CLI still supplies separate classification and presentation overrides in the SBOM writer. Twelve parent/current refusal comparisons across all three builders preserve CLI risk 1 and identical component properties/types; raw direct-API risk 0 occurs in both versions, while emitted finding paths retain risk 1. The committed test_cli_mlflow_sbom_classifies_source_before_display_bound and literal-input/emitted-path API regressions exercise these distinctions; all 13 MLflow SBOM cases passed. The original refusal builders had no acquisition file metadata to preserve. Mapping raw API inputs to emitted identities would change this baseline contract, so no additional mapping is needed.
A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review
Once complete, Codex will leave suggestions, or a comment if no findings are found.
A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review
Once complete, Codex will leave suggestions, or a comment if no findings are found.
The reason will be displayed to describe this comment to others. Learn more.
Honor the retry error sanitizer
When a caller supplies sanitize_error, this debug log still renders the raw exception because neither this wrapper nor the decorator factory invokes the callback. In verbose/debug environments, a credential-bearing provider exception or signed URL can therefore be written to logs despite the caller explicitly supplying a masking policy; the still-accepted parameter has silently become a no-op. Apply the callback before terminal formatting and continue forwarding it through retry_with_backoff and retry_cloud_operation.
The reason will be displayed to describe this comment to others. Learn more.
The callback's masking semantics are intentionally retired under the raw local evidence policy. The compatibility comment records why the keyword remains accepted. The retained tests check all three entry-point signatures, retry attempts, original exception identity, and raw single-line diagnostics. A failure-path probe also passed a masking callback to each entry point: each retried twice, retained the original error, made zero callback calls, and escaped the synthetic newline. This is keyword compatibility, not preservation of the former masking callback behavior.
The reason will be displayed to describe this comment to others. Learn more.
Keep signed URLs out of cache metadata
When the default cloud cache is used with a presigned or userinfo-bearing URL, this value is persisted verbatim in cache_metadata.json by cache_file(), leaving reusable credentials on disk after the scan. The SHA-256 field already provides the query-sensitive cache identity and url_display is not read anywhere, so retain a credential-redacted display value rather than the raw URL; restrictive file permissions do not protect copies in backups or diagnostic bundles.
The reason will be displayed to describe this comment to others. Learn more.
Raw url_display is intentional local cache metadata under the raw local evidence policy; the policy is explicit in commit 77b8963ad195f3ba228ad2566e63e2abd2a9e60e. URL metadata retains query-sensitive SHA-256 identity, while metadata permissions remain private on POSIX. Existing cache tests verify signed-URL retention, identity, permissions, and migration of legacy entries. The credential-retention observation is correct; restoring masking would reverse the agreed behavior.
A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review
Once complete, Codex will leave suggestions, or a comment if no findings are found.
Preserve source type when serializing oversized identifiers
modelaudit/integrations/sbom_generator.py:149
An oversized stream or Hugging Face classification path does not survive a saved-result round trip. serialize_source_value() replaces a source identity over 256 KiB with a modelaudit-source:...<sha256> identifier; returning that identifier here means _get_component_type() can no longer see the original suffix. For example, a long URL ending in model.pkl is classified as a model in memory but as a generic file after JSON reload. Persist a bounded classification suffix/type separately from the emitted identifier and use it here.
Please review the current head after integration with main 2a5185a. The new changes preserve the upstream SafeTensors FDICT fix and deterministic cache cleanup regressions. Existing review-thread replies document validated fixes, parent comparisons, and the explicit raw local evidence policy in AGENTS.md. Please verify current source and avoid repeating already addressed or intentional behavior.
A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review
Once complete, Codex will leave suggestions, or a comment if no findings are found.
The reason will be displayed to describe this comment to others. Learn more.
Documented the legacy field explicitly in the public report contract: redacted_value retains its name for compatibility, but now contains bounded raw evidence and provides no masking guarantee. The serialized key and detection behavior are preserved. The focused detector/text/passive-sidecar cohort passes all 96 cases, including raw Basic credentials and bounded previews.
The reason will be displayed to describe this comment to others. Learn more.
The README now explicitly documents the same legacy redacted_value contract for detector and passive-sidecar findings: retained key, bounded raw evidence, no masking guarantee. Both producers keep that schema. The committed passive Basic-auth finding and bounded-preview tests pass as part of the 96-case focused cohort; no runtime rename or alias is needed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scan reports and diagnostics now retain raw source identifiers and evidence, including embedded credentials. Remove output masking from CLI, SARIF/SBOM, CatBoost and selected detector previews while retaining evidence bounds, detections and failure classification.
This is PR 3 of the six-PR simplification stack, now targeting
mainafter PRs 1 and 2 landed. Normalization remains only where it determines finding identity, grouping or classification. Stream, directory-owner, Hugging Face and MLflow producers retain those inputs separately from raw fields, including across aggregation and saved JSON. Direct SARIF and SBOM APIs retain their original classification behavior. Producer metadata carries the original SBOM type/risk inputs through emitted raw source keys and saved JSON; the CLI retains its historical source classification.file_metadatanow uses emitted source keys: ordinary sources stay raw, and oversized source identifiers use bounded preview/digest identifiers shared by keys and references. Consumers must use those keys; previously masked keys are no longer lookup aliases. Metadata values and attribution among distinct sources are preserved. README and changelog document the intentional output/API change. The exportedredact_huggingface_url_for_displayandredact_huggingface_urls_in_texthelpers are retired; callers that need masking must apply their own presentation policy. Debug help and output tell users to inspect raw diagnostics before sharing.Validation:
Current-main integration (
2a5185a): preserves the upstream SafeTensors FDICT fix and deterministic initial/retried cache interruption fixture. Focused cache/routing checks and five neighboring compression cases pass on this branch; real benign/malicious CLI scans preserve scanner selection, JSON findings, SHA-256 metadata, and exit codes 0/1. Ruff checks pass. Changed-file mypy passes; full local mypy retains the previously documented optional TensorFlow baseline error. Hosted CI validates the published head.Both interruption cleanup tests control initial/retried capture deterministically and verify that every acquired monitor closes. Retained-interruption cases pass on each current prefix; leak mutations are rejected.
The latest base integration retains the reviewed scanner/test changes and current released dependency floor. A fresh-process logging test fixes deterministic CLI-module pollution exposed by Windows CI; all 33 ordered logging/debug/interruption cases pass on each updated prefix. The released dependency and evidence/debug cohort also passes all 368 cases per prefix. The upstream cache retry regression is included unchanged.
Final review follow-up moves report implementation details into contributor documentation and annotates the renamed JFrog regression. The earlier AST audit confirmed all 247 then-changed/new test functions were typed; 13 JFrog CLI regressions pass on each corrected prefix. Seven paired actual CLI cases preserve parent behavior for malformed URLs and bracketed local paths.
Follow-up validation: 347 focused tests pass on each corrected PR3–6 prefix. All 22,852 collected root-suite node IDs fit the Windows environment-variable limit after assigning short IDs to the two large-input tests; their inputs/assertions are unchanged. Real debug CLI, JSON and help invocations pass. Twenty-four baseline/candidate Hugging Face exception-classification and inventory cases match.
Before the portability/documentation follow-up, the committed prefix passed 1,948 affected tests, with three Windows-only and two opt-in integration skips. Import and native-artifact checks confirmed this checkout. Full Ruff format/lint passed; mypy checked 495 files and reports only the existing unreachable optional TensorFlow import at
tests/scanners/test_weight_distribution_scanner.py:2122.Parent/candidate comparisons cover 150 MLflow failure cases, 283 stream identities, and eight real JAX/TensorFlow directory-owner variants. Independent public-path probes verify malicious pickle stream aggregation, actual MLflow/Hugging Face CLI failures, saved-result identities, rule tags, grouping/counts and SBOM risk attribution.
333 SBOM comparisons preserve the distinction between literal public API inputs and CLI source normalization. Independent checks cover 28 direct type/size/hash cases and 36 CLI path-tracking cases, plus distinct signed sources with different content hashes.
Seventeen additional end-to-end cases pass unchanged on the parent and candidate, covering public emitted-asset/metadata exports, stream failures, empty/incomplete scans, Hugging Face errors and actual CLI auto-streaming. Historical risk grouping of query variants remains intact.
Eighteen final regressions also pass unchanged on the parent: long MLflow sources, bounded display collisions, literal versus emitted-path SBOM exports, and malformed-Unicode SARIF paths. Classification precedes display truncation.
Bounded MLflow locations remain distinct through saved JSON, including repeated/reordered inputs and per-source risk attribution. Stream, cloud, JFrog and retry terminal outputs filter control characters while preserving raw result/error evidence. New parent comparisons cover the saved-result and terminal contracts, including cache rejection warnings and JFrog debug output. Hugging Face probe diagnostics also retain terminal formatting.
Thirty-three parent-derived regression cases cover oversized-source serialization joins, literal alias collisions, and Hugging Face failures. Source-owned exceptions preserve their original classification, type, cause, and string-conversion count while reporting raw evidence. Direct and saved-result SBOM exports agree.
Terminal diagnostics escape CR/LF/tab as well as other unsafe controls; evidence previews retain their existing multiline contract. Parent/candidate probes cover quoted MLflow credentials and diagnostic output without changing acquired sources.
Oversized finding text retains its historical bound and SARIF fingerprint through JSON reload; an actual TFLite custom-operator scan verifies the original fingerprint. Shared source IDs preserve artifact attribution through direct SARIF and saved JSON, including literal identifiers that collide after path or URI normalization. Bounded previews escape unpaired Unicode characters while hashing the original source. JSON export remains available when the working directory has been removed or is inaccessible; Artificial source identifiers are explicitly relative, and conservative basename reservation keeps saved sources distinct across directory changes; actual SARIF rendering retains its original errors. Nested evidence retains its original depth budget. Hugging Face dry-run metadata failures keep the original failure category.
Earlier paired CLI runs cover clean, malicious, corrupt and missing inputs with exit codes 0/1/2. The final correction adds real producer/CLI regressions; no assertions were weakened to accept changed detection or fail-closed behavior.
Local full-suite, cross-platform and live remote acquisition completion is not claimed. One live Hugging Face inventory test requires unavailable SOCKS support (
socksio); the same failure reproduces on the unchanged parent. External transport is mocked in deterministic remote-source tests. Upstream-pinned mypy 2.4 could not be installed through this devbox's index; local checks use 2.3.1 with the installed Python 3.12 profile.Measured reduction: 9,622 maintained lines — production −4,352; tests −5,306; documentation +36. Counts include all retained identity helpers and added regression tests. Generated code is unchanged.
PRs 1: tooling and 2: tests have landed. Remaining review order: 3: raw evidence → 4: picklescan → 5: scanning/results → 6: acquisition/cache/progress. PR3 targets
main; PRs 4–6 target the preceding branch.The canonical agent guide explicitly preserves raw local evidence and prohibits reintroducing credential redaction, while retaining terminal escaping, permissions, and detections.
The README explicitly documents that the legacy
redacted_valuefield remains for compatibility and contains bounded raw evidence without masking guarantees. Detector and passive-sidecar schemas remain consistent; all 96 focused legacy-field cases pass.The security policy, canonical agent guide, user security model, and CVE triage guide explicitly define unredacted local output as supported behavior and accepted risk. Missing masking is not a vulnerability by itself and must not be repaired by reintroducing redaction. This includes secrets and credential-bearing URLs in local results, reports, logs, and cache metadata. Unauthorized host-data access or transmission remains in scope; telemetry collection limits are unchanged. This policy follow-up changes documentation only and passes formatting/link checks.