Repository navigation
Replies: 2 comments
|
You should use enabled_precisions as follows model = torch.compile(model, fullgraph=True,
backend="tensorrt",
options={
"enabled_precisions": {torch.float16},
},) |
0 replies
|
autocast isn't really doing anything here, it just controls precision in eager pytorch mode and that info never reaches the tensorrt builder. by default the builder is set up for fp32, so when the graph ends up with fp16 layers from autocast it complains. you need to explicitly pass enabled_precisions in options: import torch
import torch_tensorrt
with torch.no_grad():
model = torch.compile(
model,
backend="tensorrt",
options={"enabled_precisions": {torch.float16}},
fullgraph=True,
)
outputs = model(inputs)you can drop autocast entirely, not needed. if you want a mix of fp32/fp16 just put both in the set: options={"enabled_precisions": {torch.float32, torch.float16}}also worth noting, if your input tensors are already fp16, that's separate from enabled_precisions, the latter just lets tensorrt use fp16 kernels internally, doesn't set the input dtype. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
When trying to use fp16, I get the following error:
I'm configuring this as follows:
All reactions