Repository navigation
Add free GPU memory limits to runtime profiles - #2679
Baiju Meswani (baijumeswani) merged 2 commits into
Conversation
Match optional inclusive free-memory bounds alongside total-memory and integrated-device conditions, reject unreachable or ambiguous profiles in C++ and Python, and document snapshot limitations. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 56a8e51c-d65c-4713-bd43-4d24288ba5b1
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The C++ and Python implementations agree, handle physical memory constraints correctly, and have focused coverage.
Review effort: Balanced
Findings: None
What changed in this PR
Adds free GPU memory constraints to runtime-profile selection while preserving legacy behavior.
Changes:
- Adds inclusive free-memory bounds and physical-range validation.
- Applies CUDA-reported free memory during profile selection.
- Updates ABI version, documentation, and focused tests.
| File | Description |
|---|---|
src/config.h |
Adds free-memory eligibility fields and device facts. |
src/config.cpp |
Parses, validates, and matches free-memory bounds. |
src/models/runtime_profiles.cpp |
Passes CUDA free memory into profile selection. |
src/smartptrs.h |
Bumps the add-on interface version. |
src/python/py/models/builder_config.py |
Mirrors C++ validation and overlap rules. |
test/cpp/engine/config_tests.cpp |
Tests matching, validation, and physical constraints. |
test/cpp/engine/engine_test_doubles.h |
Supports configurable free memory in tests. |
test/python/builder/test_builder_config.py |
Tests Python validation and overlap detection. |
docs/model_package.md |
Documents runtime semantics and constraints. |
docs/ModelBuilderConfiguration.md |
Documents builder validation behavior. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Tianlei Wu (tianleiwu)
left a comment
There was a problem hiding this comment.
Overall this looks correct and well tested. The inclusive bounds, the "free <= total" feasibility check in both the C++ EligibilitiesIntersect and the Python overlap check, and the C++/Python validation parity all check out. I traced the boundary cases in the new tests (e.g. free 80/81 and total max 150 vs. free min 160) and they behave as the tests expect. Free memory comes from cudaMemGetInfo on the same device used for total memory, and kDeviceInterfaceVersion is the only place the add-on handshake is checked, so the bump to 9 is sufficient.
No blocking issues. Two small non-blocking suggestions are inline:
- Selection is now environment-dependent, but nothing reports which facts were used.
- A short note on the
std::nulloptfree-memory semantics.
Minor test-coverage idea: add one case combining is_integrated with free-memory bounds. The current overlap tests cover free bounds and the integrated flag separately, and the combined AND-of-three-conditions path in EligibilitiesIntersect / MatchesEligibility is not exercised together.
Report selected and base profile choices with their observed CUDA memory and integration status when logging is enabled; document unknown free-memory matching and cover combined eligibility. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 56a8e51c-d65c-4713-bd43-4d24288ba5b1
…o-speech-session-device Main moved kDeviceInterfaceVersion to 9 (microsoft#2678, microsoft#2679) while this branch had taken 8 for the appended IsHostAccessible(). The merged DeviceInterface has both main's GetIsIntegrated() and IsHostAccessible(), a layout neither 8 nor 9 describes, so the version is now 10.
Why
Runtime profiles can already use total GPU memory and whether the GPU is integrated. Total memory does not tell us how much memory is currently free. A profile sized for a device's capacity may not be a good choice when other work is using that memory.
What changed
minimum_free_device_memory_bytesandmaximum_free_device_memory_bytes. Both bounds are inclusive. They work alongside the existing total-memory bounds andis_integratedcondition.Checks
Free GPU memory is a snapshot, not a reservation. On systems where the CPU and GPU share memory, it is not a measure of free OS memory. Profile authors should still choose conservative settings and check host-memory headroom.