What is an AI/LLM penetration test?+
An AI/LLM penetration test is a structured security assessment of large language model systems and the infrastructure around them. Unlike traditional pentests that target code and servers, an AI pentest evaluates prompt-handling logic, guardrails, tool calling, retrieval pipelines, memory, model outputs, and the full application stack that exposes the model. Berghem tests using the OWASP LLM Top 10, MITRE ATLAS, and our proprietary Promptware Kill Chain, attempting prompt injection, jailbreaks, data leakage, model theft, unsafe tool invocation, and adversarial inputs. The result is a clear picture of real-world AI risk with reproducible proofs of exploitation.
What is the Promptware Kill Chain?+
The Promptware Kill Chain is a 7-stage attack framework based on the research by Schneier et al. (2025) and adapted by Berghem for practical offensive use, modeling how adversaries compromise AI systems — from initial access through prompt injection or malicious plugins, to privilege escalation, reconnaissance, persistence, command and control, lateral movement, and actions on objectives. We validated the framework after documenting 21 real-world promptware incidents in 2025–2026, 15 of which showed four or more kill chain stages and 57% of which maintained active persistence. It gives defenders a shared language to describe AI attacks and Berghem a repeatable methodology for red teaming.
Is OWASP LLM Top 10 covered?+
Yes. Every Berghem AI/LLM penetration test explicitly covers the full OWASP LLM Top 10, including prompt injection, insecure output handling, training data poisoning, model denial of service, supply chain vulnerabilities, sensitive information disclosure, insecure plugin design, excessive agency, overreliance, and model theft. We also map findings to MITRE ATLAS, NIST AI RMF, and ISO 42001 where applicable. Clients receive a coverage matrix showing which techniques were exercised, which controls held, and which require remediation — giving leadership and engineering teams an unambiguous view of AI risk against the industry standard.
How long does an AI security assessment take?+
A typical AI security assessment takes between two and six weeks, depending on scope and depth. A focused LLM penetration test against a single chatbot or RAG pipeline usually completes in 2–3 weeks. A comprehensive AI red team exercise, covering multiple agents, tools, and data sources using the full Promptware Kill Chain, can take 4–6 weeks. AI governance and assessment engagements — where we also review architecture, data flows, and regulatory posture — are scoped to the organization's size. We always agree timelines and rules of engagement with the client before work begins.
Do you handle compliance with the EU AI Act?+
Yes. Berghem helps organizations meet emerging AI regulatory obligations, including the EU AI Act, Brazil's PL 2338 AI bill, ISO/IEC 42001, and NIST AI RMF. Our AI governance practice maps your AI systems against the EU AI Act risk classification — prohibited, high-risk, limited, and minimal — and identifies the technical and organizational controls required for each category. We support risk management, documentation, transparency, human oversight, post-market monitoring, and conformity assessments. For high-risk systems, we align our offensive testing with EU AI Act Article 15 requirements on accuracy, robustness, and cybersecurity.