Of course, immediately after posting this I found the solution.
I was returning the logits (on cuda) from the runner, and moving them to cpu in the api code.
Making a custom runner where the logits are moved to cpu before being returned appears to have solved it