Building llama.cpp with CUDA

- 9 mins read

Series: Capstone

Using llama.cpp with CUDA and Testing Local Models

Howdy there everyone and welcome back to the next leg of my capstone journey. Which involves actually spinning up some locally hosted models and testing them out. Over the next while, we’re going to be deploying a few models, seeing how they run on my old RTX 2060 laptop, benchmarking their security reasoning skills and their ability to call tools from a MCP server. We’re not covering all of that right now, we’ll start with deploying our inference engine of choice (llama.cpp) and doing some hardware benchmarking just to make sure the models run fine on my hardware. So without further adieu let’s just get into this yeah?

The Test System

I started by taking an inventory instead of assuming the machine could handle the job. The system was running Ubuntu on an Intel Core i7-8750H with 14 GiB of reported system RAM, an NVIDIA GeForce RTX 2060 with 6,144 MiB of VRAM, and enough local storage for several 5–6 GB GGUF files. Pasted image 20260912160716.png

Pasted image 20260912160903.png

Pasted image 20260912160920.png

Pasted image 20260912161042.png

Pasted image 20260912162244.png My 6 GB VRAM ceiling really shapes how this all works out. I chose Q4_K_M quantizations and let llama.cpp fit layers automatically instead of pretending this laptop was a dedicated inference server as I am running a few other services in here.

Building llama.cpp with CUDA

The base prerequisites are straightforward:

sudo apt install git gcc g++ cmake nvidia-cuda-toolkit libssl-dev -y

I also used NVIDIA’s official repository instructions for the correct Ubuntu release and architecture rather than relying on a generic repository snippet. NVIDIA keeps the current steps in its CUDA Installation Guide for Linux. After installing the prerequisites, I recorded the tool versions used for the build: Pasted image 20260912170427.png Once the compiler, CMake, CUDA toolkit, and OpenSSL development headers were in place, I cloned llama.cpp:

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp

Pasted image 20260912172630.png Then I configured a CUDA-enabled release build with OpenSSL support and compiled it:

cmake -B build -DGGML_CUDA=ON -DLLAMA_OPENSSL=ON
cmake --build build --config Release

The build took a while on this machine, but it completed and produced the binaries we wanted. Pasted image 20260912172927.png

Pasted image 20260912173950.png

Pasted image 20260912212104.png I checked the CLI and server versions, then made sure llama.cpp could actually see the GPU:

./build/bin/llama-cli --version
./build/bin/llama-server --version
./build/bin/llama-cli --list-devices

Pasted image 20260912212311.png

Pasted image 20260912212629.png That last check is important. A successful compile does not necessarily mean that the runtime has found a CUDA device, so make sure to double check.

Pulling Foundation-Sec from Hugging Face

llama.cpp can resolve and download a compatible Hugging Face artifact with -hf, so the first launch was one command:

./build/bin/llama-cli \
  -hf 'fdtn-ai/Foundation-Sec-8B-Reasoning-Q4_K_M-GGUF:Q4_K_M'

Pasted image 20260912235318.png The upstream model is Foundation-Sec-8B, a cybersecurity-focused model from Foundation AI at Cisco. The tested GGUF was the Q4_K_M artifact selected by the -hf reference above. The model loaded. Our first interaction, however, went a little sideways. Pasted image 20260912235837.png

Pasted image 20260912235925.png The response appeared completely unrelated to our prompt. However, looking through tokenizer_config.json provided an important clue: the model-supplied chat template injects a detailed system message describing the model as “Metis,” constraining it to cybersecurity, and specifically discussing cloud configuration formats such as JSON and Terraform. Pasted image 20260913000933.png

Pasted image 20260913001002.png That did not prove exactly why a particular token was generated, but it did explain why the model had such a strong bias towards a cloud-security answer. It also reminded me that “the model” is not just a weights file. The tokenizer, chat template, runtime, sampling settings, and conversation state all affect the interaction. Just something to keep in mind going forward. The answer itself also had some problems. It presented JSON containing a // comment, even though JSON does not support comments. It also described a deny condition as though it restricted access to the named source rather than denying that source. On a fresh session, the model largely repeated its own system prompt and failed to answer the user prompt; on another retry, it recognized the mistake. Pasted image 20260913001946.png

Pasted image 20260913003436.png

Pasted image 20260913004105.png I then switched to a cybersecurity prompt example used by the publisher. Cisco Foundation AI’s official cookbook at revision a9165a769d32f9dfa9d581aeb93110411a7da34a records CWE-117 for its CVE-2015-10011 demonstration. The Reasoning model card reuses that CVE as a chat example but does not publish an expected response. Pasted image 20260913004740.png

Pasted image 20260913004807.png That comparison needs a large asterisk: the cookbook and my llama.cpp run differed in model variant, prompt formatting, generation length, and runtime. It is evidence of the publisher’s intended classification for the example, not a controlled reproduction or a broad measure of model accuracy.

A hardware feasibility test

After confirming that inference worked, I moved to a repeatable resource screen. The goal was not to rank intelligence, safety, or security performance. It was to see whether each quantized model could load and sustain generation on this specific laptop under some common settings. For Foundation-Sec, the launch command was:

./build/bin/llama-cli \
  -hf 'fdtn-ai/Foundation-Sec-8B-Reasoning-Q4_K_M-GGUF:Q4_K_M' \
  -c 4096 \
  -n 512 \
  -t 6 \
  -tb 6 \
  -ngl auto \
  -fit on \
  --temp 0.3 \
  --seed 42 \
  --multiline-input

That corresponds to a 4,096-token context allocation, 512-token output limit, six generation threads, six prompt-processing threads, automatic GPU placement and fitting, temperature 0.3, seed 42, the model-provided chat template, one parallel sequence, F16 K/V cache, and automatic Flash Attention. The compiled defaults were a logical batch size of 2,048 and a physical microbatch size of 512. b Pasted image 20260913023415.png To capture GPU telemetry every second, I used a separate terminal:

mkdir -p ~/capstone/evidence/foundation-sec/resource-generation/run-01
nvidia-smi dmon -s pucm -d 1 \
  | tee ~/capstone/evidence/foundation-sec/resource-generation/run-01/gpu-dmon.txt

The pucm groups capture power and temperature, utilization, clocks, and framebuffer/BAR1 memory. In a third terminal, I recorded host activity:

vmstat 1 \
  | tee ~/capstone/evidence/foundation-sec/resource-generation/run-01/system-vmstat.txt

I let both collectors run for five seconds before generation, through the response, and for five seconds afterward. Each model received the same hypothetical SOC scenario: failed SSH logins followed by a successful login, sudo, a script download, and a new process. The prompt requested evidence collection, benign explanations, compromise indicators, relevant ATT&CK techniques, containment, and escalation criteria while explicitly prohibiting invented evidence. Now while it was good to get a little insight into these models security reasoning capabilities, that is not really the point of this test run. This is just to make sure that these models can perform on my hardware setup, although a proper security reasoning test will be happening soon! Pasted image 20260913024447.png

Results

All four Q4_K_M models passed this resource-feasibility screen on the RTX 2060 host. That means they loaded, generated, and completed the resource run without an observed out-of-memory condition or instability. It does not mean their answers were equally complete or correct.

Model Generation Prompt processing Active/loaded VRAM Avg. GPU power Avg. GPU SM Avg. memory controller Peak/range temperature Avg. host user CPU
Foundation-Sec-8B-Reasoning Q4_K_M 30.6 tok/s Not recorded ~4.7 GiB loaded 83.19 W 52.31% 51.38% 63–71°C 46%
Granite 4.2 8B Q4_K_M 17.9 tok/s 444.0 tok/s ~4,688 MiB 48 W ~35% 29% 61°C ~47%
Ministral 3 8B Reasoning Q4_K_M 13.1 tok/s 468.9 tok/s ~4,638 MiB 37 W 21.8% 17.3% 60°C 48.7%
Qwen 3.5 9B Q4_K_M 10.6 tok/s 232.2 tok/s ~4,692 MiB 36.1 W 18.4% 14.6% 60°C ~49.4%

A few details are worth keeping alongside the table.

Foundation-Sec-8B-Reasoning

Foundation-Sec was tested with llama.cpp build b10934-acecd5603. It generated the full 512-token allowance at 30.6 tokens per second. Loaded GPU allocation was approximately 4.7 GiB. During sustained generation, GPU SM utilization averaged 52.31%, memory utilization averaged 51.38%, power averaged 83.19 W, and temperature ranged from 63°C to 71°C. Host user-CPU utilization averaged 46%. There was no swap-in, swap-out, I/O wait, out-of-memory condition, crash, or progressive memory growth. After generation, GPU utilization returned to 0%, power returned to 2 W, and temperature returned to 48°C while the model remained loaded. The simultaneous CPU and GPU activity is consistent with hybrid inference under automatic fitting, although the compact llama-cli interface did not expose the exact number of offloaded layers.

Granite 4.2 8B

Granite was tested with llama.cpp build b10934-acecd5603. The process occupied approximately 4,688 MiB of the 6,144 MiB GPU, leaving approximately 1,456 MiB available. It generated at 17.9 tokens per second and processed the prompt at 444.0 tokens per second. During steady generation, GPU utilization averaged approximately 35%, memory-controller utilization 29%, power 48 W, and temperature peaked at 61°C. Host user CPU averaged approximately 47%. Existing swap allocation remained constant. There was one isolated 1,116 KiB swap-in interval, but no swap-out, blocked processes, or I/O wait. I observed no OOM, thermal, paging, or stability issue. The 512-token ceiling stopped the response before it produced a final investigation plan. The run is valid for resource feasibility, not for scoring SOC-task completeness.

Ministral 3 8B Reasoning

Ministral was also tested with build b10934-acecd5603. It used approximately 4,638 MiB of active GPU framebuffer, leaving 1,506 MiB available. Prompt processing reached 468.9 tokens per second and generation reached 13.1 tokens per second. During steady generation, GPU power averaged 37 W, SM utilization 21.8%, memory-controller utilization 17.3%, and temperature peaked at 60°C. Host user CPU averaged 48.7%. I observed no active swap-out, sustained I/O wait, OOM condition, or instability. The seeded output reproduced the prior response, but visible reasoning consumed the 512-token allowance before a final investigation plan appeared. Again, that prevents a fair completeness score.

Qwen 3.5 9B

Qwen was the largest artifact in the set. With build b10934-acecd5603, it occupied approximately 4,692 MiB of GPU framebuffer and left approximately 1,452 MiB available. It processed the prompt at 232.2 tokens per second and generated at 10.6 tokens per second. During steady generation, GPU power averaged 36.1 W, SM utilization 18.4%, memory-controller utilization 14.6%, and temperature peaked at 60°C. Host user CPU averaged approximately 49.4%. I observed no active swap-out, I/O wait, OOM condition, or instability. Its response exhausted the 512-token allowance before producing a final answer and included a preliminary, unsupported attribution of attack objectives. I did not score task completeness.

What this test actually tells us

The useful result is modest: a laptop-class RTX 2060 with 6 GB of VRAM can run all four of these Q4_K_M models through a CUDA-enabled llama.cpp build at usable interactive speeds. Automatic fitting kept active allocations below the card’s capacity, and the runs did not expose a thermal or memory-stability blocker. Foundation-Sec was the fastest generator in this setup, but its early behavior also supplied the clearest warning against equating throughput with reliability. Granite, Ministral, and Qwen all hit the output ceiling before delivering a final plan. That is partly a test-design issue as reasoning models can spend a large portion of a small token budget on visible reasoning, so that will need to be addressed in future tests. There are other comparability limits. This was one host, one broad prompt, one run configuration, and quantized models from different publishers or converters. Automatic fitting may not choose identical CPU/GPU placement for every architecture. Foundation-Sec’s prompt-processing rate was not recorded. The metrics support a bounded deployment claim on this machine, not a general leaderboard.

Conclusion

So good news, all of these candidate models appear to run reasonably well on my setup here. Great, that brings us to the next step, which is evaluating these models across two different areas. Security reasoning and their ability to leverage tool/function calls via MCP. After we run some tests we’ll select the model that performs the best across both and continue the rest of our experiments with that model. That’s important to note, because it may not be the best and either one, but as long as it can do both tasks fairly decently that will be good enough for this project. So, with all that being said, stay tuned for more!