active_provider: cpu in /api/health)nvidia-smi
python -c "import torch; print(torch.cuda.is_available(), torch.version.cuda)"
python -c "import onnxruntime as ort; print(ort.get_available_providers())"
If PyTorch sees CUDA but onnxruntime only shows CPUExecutionProvider, the CUDA
shared libraries aren't on the dynamic linker's path. Argus pre-loads them automatically at
startup. If it still doesn't work, check whether both onnxruntime and
onnxruntime-gpu are installed:
python -m pip list | grep onnxruntime
If both appear, uninstall and reinstall only the GPU version:
python -m pip uninstall onnxruntime onnxruntime-gpu -y
python -m pip install onnxruntime-gpu
Then restart Argus. Confirm GPU is working via GET /api/health — the
active_provider field shows the result.
libcudart.so.X: cannot open shared object file
onnxruntime-gpu can't find the CUDA runtime library. Argus's startup code
pre-loads CUDA libs from pip-installed nvidia packages automatically. If you still see this
error:
python -m pip list | grep nvidia
python -m pip install torch --index-url https://download.pytorch.org/whl/cu124
torch.version.cuda), ensure you
have the matching onnxruntime-gpu.
null in /api/health after restart
The face engine failed to load silently at startup. Check the log for a
Failed to load face model warning with a traceback. Most common causes:
buffalo_s)
Force a reload without restarting: PUT /api/models/{id}/activate.
faiss-cpu and torch (used by YOLO object detection) each bundle
their own OpenMP runtime, and loading both in one process segfaults on Apple Silicon —
typically the process dies the first time the object detector runs, while face detection alone
is fine.
Argus auto-detects macOS and sets ARGUS_DISABLE_FAISS=true, which skips faiss
entirely (it's never imported, so its OpenMP runtime never loads) and uses the numpy matching
fallback — equivalent results, fine until tens of thousands of enrolled faces. No action
needed; this is handled for you on the native run path.
If you still hit it (e.g. running uvicorn directly, bypassing python -m app),
set the flag yourself before starting:
export ARGUS_DISABLE_FAISS=true
Linux/CUDA is unaffected and keeps faiss for fast matching at scale.