We should not try and choose the best model overall, but rather choose the best model for a specific task.
Parameter Count
Number of parameters affects latency and hardware requirements as larger models are generally slowed and require more memory. It also correlates with better generalization and reasoning capabilties, as well as more memorized knowledge. A rule of thumb is
- Small models (
~ 8B): Can be considered as autility worker. Small models are good for drafting templates or checklists and summarizing small outputs - Medium models (
~ 20-30B): Can be considered as adaily driverthat is solid for summarization and extraction tasks. - Large models (
~ 70B+): Can be considered aspecialistthat excels in multi-step reasoning and planning.
Context Window Size
Maximum number of tokens a model can process. Larger context windows help process long scan outputs or multiple artifacts from different tool outputs. Context size != more data can be processed, context rot.
Cost
- Fast and cheap models are great for formatting, summarization or extraction.
- Stronger more capable models are suited for in-depth analysis, multi-step planning, and reasoning
Tool Support
Not all models perform the same tool invocation, requiring structure arguments and producing structure outputs. Stronger models should be used multiple tool calls in complex workflows. Models should be able to parse JSON and XML.
Capabilities
Structured extraction involves a model with minimal creativity. So we need a model that strictly follows instructions, can handle structured data well. And it should also be able to preserve critical details.
Source code handling should utilize code-focused models.
Limitations
Each model has a specific knowledge cutoff. Some models might also refuse offensive tasks or produce vague responses to prevent harm. We should thus focus on non-destructive guidance as well as extraction and analysis rather than exploitation tasks.
Evaluation
How we should evaluate a models performance.
Test Set
Create a working benchmark to test certain behaviors. We want smaller sized benchmarks as the goal is not to evaluate the context here. Here is such an example of benchmarks:
- Dataset 1: Portscan results
- 1-3 single-host port scans with 8-10 open ports each
- Containing different services and software versions
- Containing at least one unknown service
- Dataset 2: HTTP response headers
- 2-3 sets of HTTP response headers
- Containing software information in
ServerorX-Powered-Byheaders
- Dataset 3: Directory enumeration snippets
- 2-3
gobusterorffufoutputs - Containing different status codes
- 2-3
- Dataset 4: Mixed artifact bundle
- 4-6 different artifacts
- Containing scope information
- Containing technique constraints (such as
No DoS) - Combining port scan results, HTTP headers, and directory enumeration
- Adding stack trace error messages
For each dataset we should have the following tasks:
Extraction task: Extract information from the data in a specified output formatSummarization task: Summarize the information and prioritize targetsPlanning task: Propose next steps, vulnerability hypotheses, and a verification plan
We should also create a pre-defined standardized prompt template.
You are assisting with artifact extraction. Analyze the following Nmap scan output and extract open ports, associated services, and service versions.
RULES:
- Use ONLY the artifacts provided.
- If a detail is not explicitly present, output "UNKNOWN".
- Do NOT invent endpoints, versions, technologies, or vulnerabilities.
- Output MUST be a valid JSON object.
JSON SCHEMA (must match exactly):
{
"assets": [
{
"target": "",
"services": [
{
"port": "",
"proto": "",
"service": "",
"version": ""
}
]
}
],
}
ARTIFACTS:
<PASTE DATASET HERE>
Lastly, we should define a scoring system to differentiate the models performance.
| Category/Score | 0 | 1 | 2 |
|---|---|---|---|
| Factual Accuracy | Invents Facts | Minor Assumptions | No Inventions. Unknowns labeled correctly |
| Completeness | Misses major items | Misses minor items | Captures all key items |
| Format Adherence | Breaks schema | Small drift | Correct schema |
| Scope Discipline | Suggests out-of-scope items | Unclear | Fully in-scope |
| Stability | Large variance | Small variance | Consistent |
| After defining these we should run each task 3 times to ensure consistent results. |
LLM Hosting
3rd Party Cloud-Based Providers
Take care of all the setup associated with the models. Some providers allow for a broad selection of models, including proprietary ones. We can use a service that provides a lot of models like OpenRouter or Replicate or use the API of a model provider directly. Some models may have features such as tool calling without additional configuration overhead.
Locally Hosted
Runs on our local system. Restricted to only open-source models, and hardware constraints are a factor. The advantage however is strong privacy and no dependency on 3rd parties.
ollama is the go-to software for running local LLMs.
# install
curl -fsSL https://ollama.com/install.sh | sh
# start the server
ollama serve
# pull a model
ollama pull llama3.2:1b
# ollama exposes an API which is can be queried
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2:1b",
"prompt": "Hello World!",
"stream": false
}' | jq .
Private Cloud Inference
Self-hosting a model in an organization’s private dedicated infrastructure. Guarantees privacy and such. These models will typically run on high-performance cloud hardware, enables hosting large LLMs that couldn’t be done locally.
Data privacy is a major concern when handling sensitive data in an engagement, typically ruling out any 3rd party providers.
For self-hosted hosted models we should ensure our supply chain is indeed intact. In practice, many testers use multiple models for different operations.