Skip to content

test: consolidate fixtures and regression harnesses - #1869

Merged
mldangelo-oai merged 3 commits into
mainfrom
mdangelo/codex/modelaudit-02-tests
Oct 3, 2026
Merged

mldangelo-oai merged 3 commits into
mainfrom
mdangelo/codex/modelaudit-02-tests

Conversation

@mldangelo-oai

@mldangelo-oai mldangelo-oai commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Shared model/container payload builders, transport mocks, subprocess checks and assertion harnesses replace repeated test setup across the root and standalone suites. Parameterization preserves the original scenarios; only four tests disappear, each an exact duplicate of a retained body. This removes 11,738 maintained test lines, including every new helper.

The standalone fixture module remains stdlib-only. Its root test bridge changes no published package dependency. Existing masking/output expectations and tests of private runtime seams remain until the corresponding later PR changes that behavior or implementation.

Validation:

  • Collected 23,717 tests versus 23,721 originally; mapped every changed node ID and proved the four duplicate bodies and their surviving assertions equal.
  • Focused paired execution: 5,512 candidate cases and 5,516 original cases passed, with the same one existing Python-version skip in each tree. Every selected node is accounted for; oversized initial runs were completed in bounded batches.
  • All 62 explicitly selected mocked JFrog/redirect/MLflow/DVC integration-marked cases passed in each tree.
  • Ruff and formatting pass; mypy reports the same existing unreachable TensorFlow import as the baseline.
  • Review follow-up: temporary 7z fixtures use pytest-owned paths and the two modified framework tests have explicit annotations; all 164 affected cases passed (two expected optional-dependency skips).
  • Deliberately failing shared assertions retain pytest's actual/expected diagnostics.

This is PR 2 of the six-PR simplification stack. Its base is the tooling PR. Full-suite and other-platform validation are not claimed.

Review order: 1: tooling → 2: tests → 3: raw evidence → 4: picklescan → 5: scanning/results → 6: acquisition/cache/progress. Each PR targets the previous branch; PR1 targets main.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-03T06:47:44.882230Z 9423827 Manual request
🔒 Security Review ✅ Completed 2026-10-03T06:46:09.206924Z 9423827 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Workflow run and artifacts

Performance Benchmarks

Compared 13 shared benchmarks with a regression threshold of 15%.
Status: 0 regressions, 0 improved, 13 stable, 0 new, 0 missing.
Aggregate shared-benchmark median: 4.249s -> 4.300s (+1.2%).

Workload Benchmark Target Size Files Baseline Current Change Status
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_base64] nested_base64 98 B 1 266.4us 287.7us +8.0% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_hex] nested_hex 130 B 1 270.9us 281.8us +4.0% stable
padded-multi-stream-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_padded_multi_stream_upload multi_stream_padded 4.1 KiB 1 315.7us 308.2us -2.4% stable
direct-malicious-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_direct_malicious_upload malicious_reduce 52 B 1 189.7us 193.6us +2.0% stable
mixed-model-repository tests/benchmarks/test_scan_benchmarks.py::test_scan_release_candidate_repository release-candidate 547.3 KiB 32 645.56ms 657.71ms +1.9% stable
rejected-basic-auth-candidates tests/benchmarks/test_scan_benchmarks.py::test_rejected_basic_auth_candidates_scan_linearly - 371.1 KiB 1 2.403s 2.448s +1.9% stable
chunked-upload-stream tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_chunked_upload_stream chunked_stream 278.2 KiB 1 117.56ms 116.17ms -1.2% stable
warm-cache-rescan tests/benchmarks/test_scan_benchmarks.py::test_scan_warm_cached_repository_rescan release-candidate 547.3 KiB 32 143.09ms 144.50ms +1.0% stable
duplicate-heavy-registry tests/benchmarks/test_scan_benchmarks.py::test_scan_duplicate_registry_snapshot registry-snapshot 915.2 KiB 13 589.45ms 584.54ms -0.8% stable
clean-training-checkpoint tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_clean_training_checkpoint safe_large 278.2 KiB 1 114.55ms 114.01ms -0.5% stable
single-checkpoint-preflight tests/benchmarks/test_scan_benchmarks.py::test_scan_single_checkpoint_before_load single_checkpoint.pkl 183.0 KiB 1 107.79ms 107.44ms -0.3% stable
suspicious-pickle-intake tests/benchmarks/test_scan_benchmarks.py::test_scan_suspicious_pickle_intake suspicious-intake 183.8 KiB 4 126.85ms 126.48ms -0.3% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_raw] nested_raw 78 B 1 239.7us 240.1us +0.2% stable

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6649a9644c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/scanners/test_sevenzip_scanner.py Outdated
Comment on lines +67 to +68
with tempfile.NamedTemporaryFile(suffix=".7z", delete=False) as f:
temp_path = f.name

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep the 7z fixture under tmp_path

Replace the new module-level NamedTemporaryFile fixture with a tmp_path: Path-based fixture. As written, every consumer creates a file in the host-global temporary directory and depends on manual teardown, so an interrupted worker can leak artifacts and the tests are no longer self-contained as required by the repository's test guardrails.

AGENTS.md reference: AGENTS.md:L122-L124

Useful? React with 👍 / 👎.

context = {"lambda_code": "lambda x: x * 0.5", "layer_definition": True}
result = knowledge_base._validate_context(pattern, context)
assert result is True
_assert_lambda_context_validation(knowledge_base, "low", "lambda x: x * 0.5", True)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Add annotations to the modified pytest tests

Annotate the enclosing test with -> None and type its knowledge_base fixture; the same omission remains across several tests modified by this consolidation. The repository explicitly applies its typed-pytest requirement to every modified test, so changing only the body still requires updating the test signature.

AGENTS.md reference: AGENTS.md:L122-L123

Useful? React with 👍 / 👎.

@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 9423827a74

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 🎉

Reviewed commit: 9423827a74

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The extensive cross-suite test refactor lacks a claimed full-suite and cross-platform validation run.

Review effort: Balanced
Findings: None

What changed in this PR

Consolidates duplicated test fixtures, payload builders, mocks, subprocess checks, and assertions across the root and standalone suites without changing production behavior.

Changes:

  • Adds shared test helper modules.
  • Replaces repeated test setup and assertions with reusable harnesses.
  • Removes duplicate test bodies while retaining scenario coverage.
File Description
tests/​utils/​test_result_conversion.py Shares status assertions.
tests/​utils/​helpers/​test_secure_hasher.py Shares hash metadata checks.
tests/​utils/​file/​test_streaming_analysis.py Consolidates short-read checks.
tests/​utils/​file/​test_file_filter.py Reuses fixture builders.
tests/​utils/​file/​test_advanced_file_handler.py Consolidates scanner fixtures.
tests/​test_why_explanations.py Reuses malicious payload.
tests/​test_weak_hash_detection.py Shares hash detection assertions.
tests/​test_tensorflow_lambda_detection.py Shares TensorFlow detection.
tests/​test_telemetry.py Consolidates telemetry policies.
tests/​test_scanner_selection.py Reuses command payload.
tests/​test_regular_scan_hash.py Shares payload and deferral helpers.
tests/​test_pytorch_zip_detection.py Reuses malicious payloads.
tests/​test_pickle_context_filtering.py Reuses command payloads.
tests/​test_os_subprocess_detection.py Shares embedded-code scanning.
tests/​test_nightly_prerequisites.py Reuses benchmark payload.
tests/​test_lazy_loading.py Shares subprocess checks.
tests/​test_huggingface_extensions.py Shares subprocess checks.
tests/​test_false_positive_fixes.py Reuses malicious payload.
tests/​test_dill_joblib_enhanced.py Consolidates Joblib fixtures.
tests/​test_cve_2025_10155_bin_pickle.py Shares symbol assertions.
tests/​test_core_asset_extraction.py Reuses executable payload.
tests/​test_cloud_url_detection.py Shares severity assertions.
tests/​test_cache_cli.py Consolidates CLI checks.
tests/​test_auth_config.py Shares transport mocks.
tests/​scanners/​test_xgboost_scanner.py Reuses binary builders.
tests/​scanners/​test_torch7_scanner.py Consolidates scanner assertions.
tests/​scanners/​test_tflite_scanner.py Shares routing assertions.
tests/​scanners/​test_tf_savedmodel_scanner.py Reuses TensorFlow fixtures.
tests/​scanners/​test_tf_metagraph_scanner.py Shares read-failure mock.
tests/​scanners/​test_tensorrt_scanner.py Consolidates marker checks.
tests/​scanners/​test_skops_content_analysis.py Shares CVE assertions.
tests/​scanners/​test_scanner_registry.py Consolidates routing fixtures.
tests/​scanners/​test_rknn_scanner.py Reuses binary writer.
tests/​scanners/​test_r_serialized_scanner.py Shares credential assertions.
tests/​scanners/​test_pytorch_binary_scanner.py Consolidates signature checks.
tests/​scanners/​test_paddle_scanner.py Reuses boundary writer.
tests/​scanners/​test_openvino_scanner.py Consolidates result checks.
tests/​scanners/​test_numpy_scanner.py Reuses executable payload.
tests/​scanners/​test_metadata_scanner.py Shares clean-input harness.
tests/​scanners/​test_llamafile_scanner.py Consolidates runtime-risk checks.
tests/​scanners/​test_lightgbm_scanner.py Reuses cache helpers.
tests/​scanners/​test_keras_utils.py Shares text instrumentation.
tests/​scanners/​test_joblib_scanner.py Shares close tracking.
tests/​scanners/​test_coreml_scanner.py Reuses protobuf encoders.
tests/​scanners/​test_compressed_scanner.py Reuses evaluation payload.
tests/​scanners/​test_cntk_scanner.py Reuses uncached scan helper.
tests/​scanners/​test_base_scanner.py Shares failure callback.
tests/​integrations/​test_mlflow_integration.py Consolidates URI mocks.
tests/​integrations/​test_jfrog_redirect_security.py Shares response and redirect fixtures.
tests/​helpers/​text.py Adds text instrumentation.
tests/​helpers/​tensorflow.py Adds TensorFlow fixture builders.
tests/​helpers/​scanners.py Adds scanner test utilities.
tests/​helpers/​processes.py Adds subprocess assertion helper.
tests/​helpers/​pickle_framework.py Bridges standalone fixtures.
tests/​helpers/​http.py Adds streaming response double.
tests/​helpers/​frameworks.py Centralizes TensorFlow detection.
tests/​helpers/​cache.py Adds cache/result helpers.
tests/​helpers/​assertions.py Adds substring assertions.
tests/​detectors/​test_secrets_detector.py Shares detection assertions.
tests/​detectors/​test_cve_detection.py Consolidates CVE checks.
tests/​detectors/​test_compile_eval_variants.py Shares pickle assertions.
tests/​conftest.py Removes no-op cleanup fixture.
tests/​benchmarks/​test_scan_benchmarks.py Consolidates benchmark assertions.
tests/​benchmarks/​test_picklescan_benchmarks.py Reuses command payload.
tests/​analysis/​test_framework_patterns.py Removes duplicate scenario.
tests/​analysis/​test_entropy_analyzer.py Shares entropy assertions.
tests/​analysis/​test_analysis_modules.py Consolidates analyzer checks.
packages/​modelaudit-picklescan/​tests/​test_rust_engine.py Reuses pickle encoder.
packages/​modelaudit-picklescan/​tests/​test_protocol0_line_operands.py Shares operand and report checks.
packages/​modelaudit-picklescan/​tests/​test_nested_budget_limits.py Reuses command payload.
packages/​modelaudit-picklescan/​tests/​test_known_size_streams.py Reuses command payload.
packages/​modelaudit-picklescan/​tests/​test_call_graph_tkinter.py Reuses call-graph helpers.
packages/​modelaudit-picklescan/​tests/​test_call_graph_six.py Consolidates alias cases.
packages/​modelaudit-picklescan/​tests/​test_call_graph_safe_spec_resolution.py Shares subprocess/importer fixtures.
packages/​modelaudit-picklescan/​tests/​test_call_graph_local_imports.py Reuses call-graph helpers.
packages/​modelaudit-picklescan/​tests/​test_call_graph_instance_defaults.py Consolidates provider operands.
packages/​modelaudit-picklescan/​tests/​test_call_graph_execnet.py Reuses call-graph helpers.
packages/​modelaudit-picklescan/​tests/​test_call_graph_click.py Reuses call-graph helpers.
packages/​modelaudit-picklescan/​tests/​pickle_test_helpers.py Adds standalone pickle utilities.
packages/​modelaudit-picklescan/​tests/​parity_corpus.py Reuses framework payload.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Base automatically changed from mdangelo/codex/modelaudit-01-tooling to main October 3, 2026 18:12
@mldangelo-oai
mldangelo-oai merged commit 35f2b46 into main Oct 3, 2026
33 checks passed
@mldangelo-oai
mldangelo-oai deleted the mdangelo/codex/modelaudit-02-tests branch October 3, 2026 18:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants