Does the behavior change if you let the shell run for a while before requesting a prediction? Checking for the process will tell you if it’s started, but loading models etc happens after that.
Also observe the machine’s CPU/GPU load while waiting: is there work being done?






















