We should not try and choose the best model overall, but rather choose the best model for a specific task.

Parameter Count

Number of parameters affects latency and hardware requirements as larger models are generally slowed and require more memory. It also correlates with better generalization and reasoning capabilties, as well as more memorized knowledge. A rule of thumb is

  • Small models (~ 8B): Can be considered as a utility worker. Small models are good for drafting templates or checklists and summarizing small outputs
  • Medium models (~ 20-30B): Can be considered as a daily driver that is solid for summarization and extraction tasks.
  • Large models (~ 70B+): Can be considered a specialist that excels in multi-step reasoning and planning.

Context Window Size

Maximum number of tokens a model can process. Larger context windows help process long scan outputs or multiple artifacts from different tool outputs. Context size != more data can be processed, context rot.

Cost

  • Fast and cheap models are great for formatting, summarization or extraction.
  • Stronger more capable models are suited for in-depth analysis, multi-step planning, and reasoning

Tool Support

Not all models perform the same tool invocation, requiring structure arguments and producing structure outputs. Stronger models should be used multiple tool calls in complex workflows. Models should be able to parse JSON and XML.

Capabilities

Structured extraction involves a model with minimal creativity. So we need a model that strictly follows instructions, can handle structured data well. And it should also be able to preserve critical details.

Source code handling should utilize code-focused models.

Limitations

Each model has a specific knowledge cutoff. Some models might also refuse offensive tasks or produce vague responses to prevent harm. We should thus focus on non-destructive guidance as well as extraction and analysis rather than exploitation tasks.

Evaluation

How we should evaluate a models performance.

Test Set

Create a working benchmark to test certain behaviors. We want smaller sized benchmarks as the goal is not to evaluate the context here. Here is such an example of benchmarks:

  • Dataset 1: Portscan results
    • 1-3 single-host port scans with 8-10 open ports each
    • Containing different services and software versions
    • Containing at least one unknown service
  • Dataset 2: HTTP response headers
    • 2-3 sets of HTTP response headers
    • Containing software information in Server or X-Powered-By headers
  • Dataset 3: Directory enumeration snippets
    • 2-3 gobuster or ffuf outputs
    • Containing different status codes
  • Dataset 4: Mixed artifact bundle
    • 4-6 different artifacts
    • Containing scope information
    • Containing technique constraints (such as No DoS)
    • Combining port scan results, HTTP headers, and directory enumeration
    • Adding stack trace error messages

For each dataset we should have the following tasks:

  • Extraction task: Extract information from the data in a specified output format
  • Summarization task: Summarize the information and prioritize targets
  • Planning task: Propose next steps, vulnerability hypotheses, and a verification plan

We should also create a pre-defined standardized prompt template.

You are assisting with artifact extraction. Analyze the following Nmap scan output and extract open ports, associated services, and service versions.

RULES:
- Use ONLY the artifacts provided.
- If a detail is not explicitly present, output "UNKNOWN".
- Do NOT invent endpoints, versions, technologies, or vulnerabilities.
- Output MUST be a valid JSON object.

JSON SCHEMA (must match exactly):
{
  "assets": [
    {
      "target": "",
      "services": [
        {
          "port": "",
          "proto": "",
          "service": "",
          "version": ""
        }
      ]
    }
  ],
}

ARTIFACTS:
<PASTE DATASET HERE>

Lastly, we should define a scoring system to differentiate the models performance.

Category/Score012
Factual AccuracyInvents FactsMinor AssumptionsNo Inventions. Unknowns labeled correctly
CompletenessMisses major itemsMisses minor itemsCaptures all key items
Format AdherenceBreaks schemaSmall driftCorrect schema
Scope DisciplineSuggests out-of-scope itemsUnclearFully in-scope
StabilityLarge varianceSmall varianceConsistent
After defining these we should run each task 3 times to ensure consistent results.

LLM Hosting

3rd Party Cloud-Based Providers

Take care of all the setup associated with the models. Some providers allow for a broad selection of models, including proprietary ones. We can use a service that provides a lot of models like OpenRouter or Replicate or use the API of a model provider directly. Some models may have features such as tool calling without additional configuration overhead.

Locally Hosted

Runs on our local system. Restricted to only open-source models, and hardware constraints are a factor. The advantage however is strong privacy and no dependency on 3rd parties.

ollama is the go-to software for running local LLMs.

# install
curl -fsSL https://ollama.com/install.sh | sh
 
# start the server
ollama serve
 
# pull a model
ollama pull llama3.2:1b
 
# ollama exposes an API which is can be queried
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2:1b",
  "prompt": "Hello World!",
  "stream": false
}' | jq .
 

Private Cloud Inference

Self-hosting a model in an organization’s private dedicated infrastructure. Guarantees privacy and such. These models will typically run on high-performance cloud hardware, enables hosting large LLMs that couldn’t be done locally.


Data privacy is a major concern when handling sensitive data in an engagement, typically ruling out any 3rd party providers.

For self-hosted hosted models we should ensure our supply chain is indeed intact. In practice, many testers use multiple models for different operations.