You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Bring the AMDGPU (DirectX/MIGraphX umbrella) OGA path in line with the DML path for graph capture and device teardown.
Graph capture:
Keep graph capture ON by default for AMDGPU, but add a per-model opt-out via the provider option enable_graph_capture="0" (matching DML). The MIGraphX backend relies on capture being on by default.
Force capture OFF for non-decoder encoder/joiner sub-sessions (Whisper, Marian, Nemotron, Parakeet) on the AMDGPU device only, since their control-flow graphs could be incompatible with captured-graph replay. DML's existing behavior is left unchanged.
Teardown / lifetime:
Add per-model CloseAMDGPUInterface() (from Model::~Model, mirroring CloseDmlInterface) that resets the cached device allocator + init session, destroys the interface singletons, releases the OrtEnv-shared allocators, and unregisters the umbrella EP library.
Refcount the shared interface singleton so a composite that keeps several models alive (e.g. speculative decoding's target + draft) only tears the device down on the last release.
Only release the shared allocators / unregister the EP when genai owns the registration, so a host that pre-registered the library keeps it.
Clear device-backed shared initializers before teardown so their GpuMemory frees through a live allocator.
Make the AMDGPU interface non-cached in OrtGlobals::GetDeviceInterface (like DML), so a lookup after per-model teardown rebuilds a fresh instance instead of returning a dangling pointer.
Bring the AMDGPU (DirectX/MIGraphX umbrella) OGA path in line with the DML
path for graph capture and device teardown.
Graph capture:
- Keep graph capture ON by default for AMDGPU, but add a per-model opt-out
via the provider option enable_graph_capture="0" (matching DML). The
MIGraphX backend relies on capture being on by default.
- Force capture OFF for non-decoder encoder/joiner sub-sessions (Whisper,
Marian, Nemotron, Parakeet) on the AMDGPU device only, since their
control-flow graphs are incompatible with captured-graph replay. DML's
existing behavior is left unchanged.
Teardown / lifetime:
- Add per-model CloseAMDGPUInterface() (from Model::~Model, mirroring
CloseDmlInterface) that resets the cached device allocator + init session,
destroys the interface singletons, releases the OrtEnv-shared allocators,
and unregisters the umbrella EP library.
- Refcount the shared interface singleton so a composite that keeps several
models alive (e.g. speculative decoding's target + draft) only tears the
device down on the last release.
- Only release the shared allocators / unregister the EP when genai owns the
registration, so a host that pre-registered the library keeps it.
- Clear device-backed shared initializers before teardown so their GpuMemory
frees through a live allocator.
- Make the AMDGPU interface non-cached in OrtGlobals::GetDeviceInterface (like
DML), so a lookup after per-model teardown rebuilds a fresh instance instead
of returning a dangling pointer.
- Scope the EP registration-ownership flag to the OrtEnv lifetime (move from a
file-static into OrtGlobals) so it can't leak stale ownership across
OgaShutdown or a failed Model construction and tear down a host's registration.
- Apply the effective graph-capture decision to both backend keys, so the
opt-out and non-decoder force-off also gate ep.migraphx.hip_graph_enable, not
just DirectML.
- Reject allocator reuse when a live model selected a different device, instead
of silently binding to the wrong device's allocator.
- Document the refcount's unsynchronized, serialized-lifecycle assumption
(matching the DML interface).
~Model dispatched teardown by reading p_device_->GetType(). For a DML model,
CloseDmlInterface() frees the singleton p_device_ points to, so the following
AMDGPU-path check re-read freed memory and aborted DML teardown (hit by the
DML CI tests). Snapshot the type once before any teardown and dispatch on it.
GenAI-owned AMDGPU registration leaks after model construction fails
src/generator/generators.h:224
An OrtEnv can outlive OgaShutdown() when the host retains it (ep/ryzenai/interface.cpp:123–125). If automatic AMDGPU registration succeeds but the base Model constructor throws, for example while loading shared initializers, ~Model() never releases the registration. OrtGlobals teardown then discards this flag without unregistering. After re-init, the surviving registration is treated as host-owned, so later model teardown skips its allocator and EP cleanup. Release genai-owned registrations during shutdown, including after failed construction, or preserve ownership while the environment survives.
…struction fails
ReleaseOwnedUmbrellaEp only ran from Model::~Model, so an owned AMDGPU
registration leaked when the host kept the OrtEnv past OgaShutdown or a
Model ctor threw. Also call it from ~OrtGlobals before env_.reset(), passing
env + the flag by reference so it doesn't re-lock the held g_ort_globals_mutex
and deadlock.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bring the AMDGPU (DirectX/MIGraphX umbrella) OGA path in line with the DML path for graph capture and device teardown.
Graph capture:
Teardown / lifetime: