Skip to content

Add free GPU memory limits to runtime profiles - #2679

Merged
Baiju Meswani (baijumeswani) merged 2 commits into
mainfrom
baijumeswani/free-memory-runtime-profiles
Oct 6, 2026
Merged

Baiju Meswani (baijumeswani) merged 2 commits into
mainfrom
baijumeswani/free-memory-runtime-profiles

Conversation

@baijumeswani

Copy link
Copy Markdown
Collaborator

Why

Runtime profiles can already use total GPU memory and whether the GPU is integrated. Total memory does not tell us how much memory is currently free. A profile sized for a device's capacity may not be a good choice when other work is using that memory.

What changed

  • Profiles can now set optional minimum_free_device_memory_bytes and maximum_free_device_memory_bytes. Both bounds are inclusive. They work alongside the existing total-memory bounds and is_integrated condition.
  • The runtime uses the free-memory value it already reads from CUDA when it creates a Model. Profiles without free-memory bounds keep their previous behavior; if no profile matches, the base configuration is used.
  • The C++ loader and Python builder reject invalid ranges and profiles that could match the same device. They account for the fact that free memory cannot exceed total memory.
  • The host/CUDA add-on interface version moves to 9 because the configuration layout changed. The binaries must be built from matching sources.

Checks

  • 33 focused C++ configuration tests passed.
  • 211 Python builder-config tests passed.
  • The changed C++ files passed the local clang-format check.

Free GPU memory is a snapshot, not a reservation. On systems where the CPU and GPU share memory, it is not a measure of free OS memory. Profile authors should still choose conservative settings and check host-memory headroom.

Match optional inclusive free-memory bounds alongside total-memory and integrated-device conditions, reject unreachable or ambiguous profiles in C++ and Python, and document snapshot limitations.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 56a8e51c-d65c-4713-bd43-4d24288ba5b1
Copilot AI balanced review requested due to automatic review settings October 6, 2026 07:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The C++ and Python implementations agree, handle physical memory constraints correctly, and have focused coverage.

Review effort: Balanced
Findings: None

What changed in this PR

Adds free GPU memory constraints to runtime-profile selection while preserving legacy behavior.

Changes:

  • Adds inclusive free-memory bounds and physical-range validation.
  • Applies CUDA-reported free memory during profile selection.
  • Updates ABI version, documentation, and focused tests.
File Description
src/​config.h Adds free-memory eligibility fields and device facts.
src/​config.cpp Parses, validates, and matches free-memory bounds.
src/​models/​runtime_profiles.cpp Passes CUDA free memory into profile selection.
src/​smartptrs.h Bumps the add-on interface version.
src/​python/​py/​models/​builder_config.py Mirrors C++ validation and overlap rules.
test/​cpp/​engine/​config_tests.cpp Tests matching, validation, and physical constraints.
test/​cpp/​engine/​engine_test_doubles.h Supports configurable free memory in tests.
test/​python/​builder/​test_builder_config.py Tests Python validation and overlap detection.
docs/​model_package.md Documents runtime semantics and constraints.
docs/​ModelBuilderConfiguration.md Documents builder validation behavior.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@tianleiwu Tianlei Wu (tianleiwu) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this looks correct and well tested. The inclusive bounds, the "free <= total" feasibility check in both the C++ EligibilitiesIntersect and the Python overlap check, and the C++/Python validation parity all check out. I traced the boundary cases in the new tests (e.g. free 80/81 and total max 150 vs. free min 160) and they behave as the tests expect. Free memory comes from cudaMemGetInfo on the same device used for total memory, and kDeviceInterfaceVersion is the only place the add-on handshake is checked, so the bump to 9 is sufficient.

No blocking issues. Two small non-blocking suggestions are inline:

  • Selection is now environment-dependent, but nothing reports which facts were used.
  • A short note on the std::nullopt free-memory semantics.

Minor test-coverage idea: add one case combining is_integrated with free-memory bounds. The current overlap tests cover free bounds and the integrated flag separately, and the combined AND-of-three-conditions path in EligibilitiesIntersect / MatchesEligibility is not exercised together.

Comment thread src/models/runtime_profiles.cpp
Comment thread src/config.h
Report selected and base profile choices with their observed CUDA memory and integration status when logging is enabled; document unknown free-memory matching and cover combined eligibility.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 56a8e51c-d65c-4713-bd43-4d24288ba5b1
@baijumeswani
Baiju Meswani (baijumeswani) merged commit 6804271 into main Oct 6, 2026
66 of 67 checks passed
@baijumeswani
Baiju Meswani (baijumeswani) deleted the baijumeswani/free-memory-runtime-profiles branch October 6, 2026 22:27
Yuri Khrustalev (ykhrustalev) added a commit to ykhrustalev/onnxruntime-genai that referenced this pull request Oct 7, 2026
…o-speech-session-device

Main moved kDeviceInterfaceVersion to 9 (microsoft#2678, microsoft#2679) while this branch had taken 8 for the
appended IsHostAccessible(). The merged DeviceInterface has both main's GetIsIntegrated() and
IsHostAccessible(), a layout neither 8 nor 9 describes, so the version is now 10.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants