While it cannot replace a human penetration tester, AI can server as a powerful tool for processing large amounts of data, generating insights, and streamlining decision-making throughout an engagement.

AI integrates into all stages of an engagement. When there is a large amount of data, the probability of missing important aspects increases, potentially leading a human tester to miss security configurations or vulnerabilities.

Modes of Testing

Manual Testing

A human executes tasks manually and makes all decisions. Highly adaptable since a human tester can easily pivot and react to unexpected circumstances. Manual work produces the strongest evidence, including clear reproduction steps and proof of compromise, with less of a chance of false positives.

Manual work is very time and labor intensive, making it a slow and expensive process.

Automated Testing

Predefined scripts and tools that execute certain repeatable tasks automatically. Include scanners, fuzzers, and exploitation tools. These are highly repeatable and scalable but lacks the ability to properly assess edge cases and also produce high noise.

Reduces the time require to execute repetitive tasks and ideal for baseline checks that are not particularly prone to errors.

AI-Augmented Testing

Combines the advantage of human expertise with automation by utilizing AI-driven analysis. AI can interpret, summarize and reason about data, rather than simply executing commands. This type of testing adheres to the principle of Automation executing, AI interpreting, and Humans deciding.

AI thrives at summarization of tool output to create structured notes, pattern recognition to match data-backend to common types of issues, and planning and prioritization by proposing hypotheses and suggesting next steps.

AI does however tend to invent details and also lacks environmental awareness as it has no understanding of the target environment or operational security. It must be closely supervised and verified.

Reconnaissance

Convert raw artifacts from automated enumeration tools into structure, summarized data. Sifting through noise and extracting the most relevant information.

Testing

Help interpret response patterns and suggest specific test cases based on observed functionality. Organize information, identify repeated patterns. Review source code and configurations files to identify misconfigurations and bugs.

Attack-Path Planning

AI can analyze all artifacts and suggest next steps for priv esc vectors. When analyzing large AD environments, where a lack of data is often not an issue, but the challenge lies in effectively prioritizing targets.

Reporting

AI can structure our notes or draft parts of the deliverable. LLMs can provide consistent value for text based tasks. AI can help adjust our writing style for different target audiences.

Risks

Overreliance can reduce the quality of the results. AI can only assist, not replace, human judgement. AI output should not be fact, it is just a suggestion. We need to assess the system as a whole and think about potential attack vectors beyond AI suggestions.

AI service providers puts the of the client data at risk.

Integration into workflows

  1. Pre-processing: preparing the raw data from different sources, such as tool outputs and manual enumeration to improve LLM performance.
    1. Normalizing the format by converting tool outputs and artifcatcs into readable blocks and avoiding mixing multiple outputs into a single text blob. We can use the following templates
    2. Redacting sensitive data when using 3rd party LLM service providers to reduce sensitive client information to avoid data exposure (credentials, email addresses, source code, internal-only hostnames or network information). Important to include contextual information in what is being redacted.
[ARTIFACT] Nmap scan
Target: 10.10.10.10
Tool: nmap
Content:
<SNIP>
  • Decision support: can help make decisions by providing guidance
    • Ask for multiple hypotheses to avoid fixation on a single attack vector. “Provide 3 hypotheses and how to validate each”
    • Make constraints explicit to communicate the boundaries and rules of engagement. “No DoS conditions”
    • Prioritzation accordingg to clearly defined criteria.
    • Maintain a decision log to help team understand why specific decisions were made.
    • AI is quite good at pattern recognition
      • What does this pattern suggest?
      • What security risks are commonly associated with the following setup
    • Am I missing something? Provide what we’ve done and information we know to reveal other vectors
    • Choosing the right tool
  • Post-processing by preparing notes and evidence into easily reviewable output.