Building llama.cpp with CUDA on Windows 11
A record of building llama.cpp, the go-to local LLM inference engine, from source on Windows 11 with the CUDA backend enabled for an NVIDIA GPU, through to actually running a model.
A record of building llama.cpp, the go-to local LLM inference engine, from source on Windows 11 with the CUDA backend enabled for an NVIDIA GPU, through to actually running a model.