Evaluating Models with EvidenceForge

- 12 mins read

Series: Capstone

Benchmarking Models with EvidenceForge

Alright, so previously we did some hardware benchmarking to validate that our selected models could reasonable run on my current hardware. Those tests all went well, which was good, but now we are getting into the real meat and potatoes of this project. Evaluating these models security reasoning skills. How are we going to go about doing that? Well, I’m glad you asked. Enter, EvidenceForge. EvidenceForge is an open-source project from Cisco Talos that can generate synthetic security telemetry from scenario definitions, keeps the generated evidence separate from the answer key, and gives me a much better basis for comparing local models than a few improvised chat prompts. Feel free to check out the repo for more details on the project, but we’re going to go ahead and get into it!

Getting Started with EvidenceForge

EvidenceForge lets us describe a security scenario and generate the telemetry that scenario would leave behind. Instead of asking a model to analyze three lines I made up on the spot, I can give it a compact evidence packet drawn from generated logs, keep a durable record of the generation process, and compare its report against the scenario I authored. That distinction matters. A useful benchmark needs at least three separate things:

  1. The scenario and ground truth — what was designed to happen.
  2. The observable telemetry — what an analyst or model would actually get to see.
  3. The model’s analysis — the conclusions it draws from those observations. If those layers get mixed together, the evaluation is compromised. If the model sees a resolved scenario file or another ground-truth artifact, it is no longer investigating. It is reading the answer key. EvidenceForge also makes the test repeatable. A scenario can be validated, generated with a fixed seed, evaluated, and recorded alongside the exact repository revision and scenario hash. That does not magically make every modeling choice objective, but it does make our experiment auditable.

Installing EvidenceForge

The project uses uv, Astral’s Python package and project manager. Start by cloning the repository:

git clone https://github.com/Cisco-Talos/EvidenceForge.git
cd EvidenceForge

Pasted image 20260917022955.png If uv is not already installed, install it and confirm that it is available:

curl -LsSf https://astral.sh/uv/install.sh | sh

uv --version

Pasted image 20260917023307.png Then let uv create the environment and synchronize the project dependencies:

uv sync

Pasted image 20260917023444.png The repository contains the code, documentation, skills for supported coding agents, and the scenario definitions that drive generation. Pasted image 20260917235554.png The directory I care about for this benchmark is scenarios: Pasted image 20260917235736.png

EvidenceForge includes example material, but I authored a compact scenario for this test because a larger prebuilt scenario would produce more telemetry than I could reasonably present to a model with a 4,096-token context window. From the repository root, I created a directory for it:

mkdir -p scenarios/compact-web-compromise

The scenario itself is a YAML file placed in that directory. Which I’m not including here as it’s kinda lengthy. The important methodological point is that the scenario definition is an input to generation and later becomes part of the benchmark’s protected ground truth.

Validate first, then preserve provenance

A scenario that happens to generate output is not necessarily a scenario I should trust. Validation catches structural problems before I spend time building the rest of the run around it. I also wanted every meaningful command to leave a durable artifact. The working directory below separates this run by model, test mode, scenario, and seed:

mkdir -p \
  bundles \
  ~/capstone/evidence/foundation-sec/model-only/compact-web-compromise/seed-42

Next, I captured validation output instead of treating it as disposable terminal text:

set -o pipefail
uv run eforge validate \
  scenarios/compact-web-compromise/scenario.yaml \
  2>&1 | tee \
  ~/capstone/evidence/foundation-sec/model-only/compact-web-compromise/seed-42/validation.txt

Pasted image 20260918011603.png set -o pipefail is worth keeping here. Without it, a failure on the left side of the pipe can be obscured by a successful tee on the right.

Deterministic generation

For the generated bundle, I fixed the random seed at 42 and named the output directory accordingly:

uv run eforge generate \
  scenarios/compact-web-compromise/scenario.yaml \
  --output bundles/compact-web-compromise-seed-42 \
  --seed 42 \
  --target default \
  2>&1 | tee \
  ~/capstone/evidence/foundation-sec/model-only/compact-web-compromise/seed-42/generation.txt

Pasted image 20260919190357.png A fixed seed is not the same thing as universal reproducibility. The repository revision, tool version, scenario contents, target, and environment still matter. That is why the provenance record sits next to the generation log. Together they make this specific bundle traceable instead of just giving it a memorable folder name. I then ran EvidenceForge’s evaluation step and captured that output too:

uv run eforge eval \
  bundles/compact-web-compromise-seed-42 \
  2>&1 | tee \
  ~/capstone/evidence/foundation-sec/model-only/compact-web-compromise/seed-42/evidenceforge-evaluation.txt

Pasted image 20260919190707.png At this point I had a generated bundle with telemetry and evaluation artifacts. I did not have a model-ready prompt yet.

Keep the answer key away from the model

The generated bundle contains more than the evidence an analyst should see. Some files exist to describe or evaluate the scenario itself. Those are useful to me as the benchmark author, but they must stay out of the model’s context. Pasted image 20260919193900.png For the model input, I worked from the generated telemetry under data rather than indiscriminately concatenating the bundle: Pasted image 20260919193959.png My rule was straightforward: if a file reveals what the scenario says happened, what the expected interpretation is, or how the run should be scored, it does not belong in the evidence packet. The model should see observables, not labels. This is especially important with a synthetic benchmark. Because I authored the scenario, it is easy for privileged information to leak into filenames, descriptions, comments, or generated metadata without feeling like an obvious answer key. Keeping generation and evaluation artifacts separate from candidate evidence reduces that risk, but the final packet still needs a manual leakage review.

Fitting evidence into 4,096 tokens

The full telemetry bundle was too large to paste into the model, and --ctx-size 4096 does not mean I can spend all 4,096 tokens on logs. The context also has to hold the system prompt, task instructions, formatting requirements, and any tokens the runtime needs within that window. Separately, the model needs enough output budget to finish its report. So I selected compact time windows around the scenario’s authored events instead of dumping every generated record. The selection process used the normalized RESOLVED_SCENARIO.yaml produced by EvidenceForge and a separate packet-plan.yaml describing which event windows to extract. A Python script then created two outputs:

Evidence packet: /home/kaleb/capstone/EvidenceForge/bundles/compact-web-compromise-seed-42/candidate-selection/candidate-evidence.txt

Audit manifest: /home/kaleb/capstone/EvidenceForge/bundles/compact-web-compromise-seed-42/candidate-selection/candidate-selection-manifest.json

Pasted image 20260920212947.png The evidence packet is what the model can inspect. The manifest is for me: it records how the candidate packet was assembled so that selection does not become an invisible, hand-edited step.

Running the models consistently

I stored the system instructions and benchmark request in separate text files and supplied them through llama-cli. Pasted image 20260920221925.png For Foundation-Sec, the command was:

./build/bin/llama-cli \
    -hf 'fdtn-ai/Foundation-Sec-8B-Reasoning-Q4_K_M-GGUF:Q4_K_M' \
    --ctx-size 4096 \
    --system-prompt-file ../prompts/system-prompt.txt \
    --file ../prompts/benchmark-prompt.txt \
    --single-turn \
    --n-predict 1600 \
    --temperature 0 \
    --seed 42 \
    --no-display-prompt \
    --verbose-prompt \
    --color off \
    --output-file ../prompts/foundation-response1.txt \
    2> ../prompts/llama.log

Pasted image 20260920223728.png For Granite 4.2, the recorded base command was:

./build/bin/llama-cli \
    -hf 'ibm-granite/granite-4.2-8b-GGUF' \
    --ctx-size 4096 \
    --system-prompt-file ../prompts/system-prompt.txt \
    --file ../prompts/benchmark-prompt.txt \
    --single-turn \
    --n-predict 1600 \
    --temperature 0 \
    --seed 42 \
    --no-display-prompt \
    --verbose-prompt \
    --color off \
    --output-file ../prompts/granite4_2-response1.txt \
    2> ../prompts/llama.log

For Ministral 3:

./build/bin/llama-cli \
    -hf 'mistralai/Ministral-3-8B-Reasoning-2512-GGUF' \
    --ctx-size 4096 \
    --system-prompt-file ../prompts/system-prompt.txt \
    --file ../prompts/benchmark-prompt.txt \
    --single-turn \
    --n-predict 1600 \
    --temperature 0 \
    --seed 42 \
    --no-display-prompt \
    --verbose-prompt \
    --color off \
    --output-file ../prompts/minstral3-response1.txt \
    2> ../prompts/llama.log

And for Qwen 3.5:

./build/bin/llama-cli \
    -hf 'bartowski/Qwen_Qwen3.5-9B-GGUF' \
    --ctx-size 4096 \
    --system-prompt-file ../prompts/system-prompt.txt \
    --file ../prompts/benchmark-prompt.txt \
    --single-turn \
    --n-predict 1600 \
    --temperature 0 \
    --seed 42 \
    --no-display-prompt \
    --verbose-prompt \
    --color off \
    --output-file ../prompts/qwen3_5-response1.txt \
    2> ../prompts/llama.log

The common settings matter. Every run used a 4,096-token context, a 1,600-token generation allowance, temperature 0, seed 42, single-turn execution, and prompt files rather than an interactive conversation. That controls several obvious sources of variation.

It does not control everything. Different model families handle internal reasoning differently, and both Granite and Qwen changed behavior when reasoning was constrained. In this hardware and token-budget regime, finishing the requested report is part of the task, we can’t have it getting cut off part way through.

What the scenario tested

The compact scenario mixed suspicious and legitimate behavior. The evidence included a web scan, shell execution by the www-data service account, routine administrator activity, and a relationship between the scan source and a domain later used as a periodic outbound TLS destination. That combination tests more than simple keyword recognition. A model has to distinguish an observation from an inference, correlate activity across sources, avoid treating every administrator command as malicious, and resist turning an opaque TLS connection into a claim about payload contents. The test was especially useful because several plausible-sounding mistakes were available:

  • calling a local socket-listing command a network scan;
  • calling an inbound administrator SSH session lateral movement;
  • treating authorized sudo iostat activity as privilege escalation;
  • treating local service-account execution as lateral movement;
  • describing periodic TLS as confirmed command and control; or
  • claiming data exfiltration without payload or directionally correct byte evidence. Here is how the models handled those traps.

Observations

The following are observations from this specific packet and run configuration. They are not general model rankings.

Foundation-Sec

Foundation-Sec found several of the important security signals. It identified the web scan, treated command execution by www-data as suspicious, and connected the Northbridge domain with the suspicious IP. It interpreted that relationship as possible command-and-control behavior, which was directionally useful. Its main problem was over-classification. It became fixated on the administrator’s activity, treated legitimate behavior as malicious, described the inbound administrator SSH session as lateral movement, and called ss -tulpn network port scanning. It also suggested exfiltration even though the packet did not establish it. In other words, Foundation-Sec was alert to the compromise story but had trouble maintaining a clean baseline. That is better than missing everything, but it would create false-positive work for an analyst.

Granite 4.2

Granite landed on the other end of the spectrum in the constrained-reasoning run. It recognized the routine administrator activity and noticed the web scan, but it was not sufficiently concerned about shell commands run by www-data. It also missed the relationship between the scanner source and the later domain resolution/TLS destination, and it ultimately classified the activity as benign. The unconstrained run showed why generation behavior belongs in the evaluation. Granite separated legitimate administration from suspicious service-account execution and caught the scan-to-destination correlation, but it spent the full generation allowance on analysis and never produced the final report. For this task, a correct thought that never becomes a usable answer is still a failed completion. Granite exposed a real accuracy-versus-completion tradeoff under a limited output budget.

Ministral 3

Ministral consistently reached the correct malicious verdict, but its tactic assignment looked driven by security keywords more than by the causal structure of the evidence. It labeled authorized sudo iostat execution as privilege escalation, local service-account execution as lateral movement, periodic TLS as confirmed exfiltration, and routine administrative SSH as possible account compromise. It sometimes mentioned benign alternatives later, but it did not reconcile those alternatives with its stronger claims. It also failed to follow the requested length and format closely enough to finish within the available output budget. The headline verdict was useful. The path it took to get there was not reliable enough for unreviewed use.

Qwen 3.5

Qwen showed the strongest evidence correlation and baseline separation in these runs. It treated iostat and the administrator’s SSH activity as likely legitimate, flagged www-data execution, and emphasized that the source associated with scanning later appeared as the server’s periodic outbound TLS destination. The unconstrained run consumed the entire generation allowance in internal reasoning and failed to deliver a final report. It also leaned too far toward confirmed command and control even though the encrypted payload was not visible. With reasoning constrained, Qwen completed the requested format and produced the strongest overall investigation in this pilot. It still made important mistakes: it misread inbound response bytes as high-volume outbound traffic, suggested exfiltration on that basis, mislabeled local discovery as lateral movement or persistence, and described suspected command and control as confirmed. So even the best qualitative result still needed an analyst to check directionality, distinguish local execution from movement between hosts, and calibrate certainty.

Conclusions

My first conclusion is that evidence correlation is a better discriminator than the final malicious-or-benign label. Several models could land on a plausible verdict. The more revealing question was whether they could explain which records belonged together without rewriting normal administration as an attack or upgrading an encrypted connection into proof of exfiltration. Second, context and output limits are part of the benchmark. On this setup, the model does not get unlimited room to think and then unlimited room to answer. Granite and Qwen both demonstrated that internal reasoning can improve analysis while simultaneously preventing completion. A security assistant that never emits the report is not useful merely because its hidden analysis was promising. Third, smaller models can provide useful investigative drafts, but none of these observations support autonomous classification. Foundation-Sec found meaningful signals but over-alerted. Granite could be too conservative or fail to finish, depending on reasoning behavior. Ministral reached the right verdict with weak causal discipline. Qwen produced the strongest completed analysis but still overstated command and control and exfiltration. My operational takeaway is narrow: a compact, carefully selected evidence packet can get useful work out of local models, especially for initial correlation and draft reporting, but the output needs evidence-level review. The model should point an analyst toward relationships worth checking, not become the source of record for what happened. Finally, provenance and leakage controls matter as much as the inference command. A fixed seed, scenario hash, tool revision, validation log, generation log, candidate-selection manifest, and saved model output turn a demo into something I can inspect and rerun. Keeping ground truth outside the prompt ensures the model is actually being tested.

Conclusion

So there we have it, that concludes the security reasoning portion of this project. In full transparency I was hoping to do a few different scenarios and see how the models did across a few different instances, but I simply do not have the time to work that in unfortunately. So instead we will be continuing to the MCP tool call benchmarking, which will be leveraging a custom MCP server I am building from scratch… so… I guess I’ll get started on that and I will see you when that’s done or something.