How to use this list
Do not read all thirty-two at a vendor in one sitting. Pick the eight or ten that matter most for your driver, ask them in a live conversation, and listen for whether the answer is specific. The pattern to watch for is simple: a provider with a documented process answers with detail and admits limits, while a provider without one answers with adjectives.
Weak answers below are not accusations of dishonesty. They are usually a salesperson repeating a phrase they have not been given the detail behind. The right response is to ask for the technical person on the next call, and to notice how quickly that happens.
If you are earlier in the process, how to buy penetration testing covers scoping and pricing, and the RFP template turns these themes into a written document.
Methodology and coverage
- Which published standards does your methodology map to, and where do you deviate from them? Good: names the Penetration Testing Execution Standard, the OWASP Web Security Testing Guide and NIST SP 800-115, then explains a deliberate deviation and the reason for it. Weak: "we follow industry best practice" with no standard named.
- Walk me through your phases on an engagement like ours. Good: a concrete sequence from pre-engagement authorisation through recon, enumeration, analysis, validation and reporting, with what your team has to do at each step. Weak: a generic four-box diagram that would fit any engagement.
- What is genuinely out of reach for your approach? Good: an honest list, usually including complex multi-step business logic, chained authorisation flaws and anything needing deep domain knowledge of your product. Weak: "nothing, we cover everything." Full coverage is not a thing that exists.
- How do you handle assets we did not know we had? Good: describes discovery, then a defined process for bringing a newly found asset into authorised scope before touching it. Weak: either no discovery at all, or discovery with no authorisation step, which is worse.
Scope and authorisation
- What authorisation do you require before testing starts, and what does it cover? Good: a signed authorisation document naming the exact assets, the window and the permitted activities, plus a requirement for third-party written permission where an asset is hosted by someone else. Weak: an email confirmation and an eagerness to start.
- What technically stops testing from touching something out of scope? Good: describes enforcement in the tooling or process, not just an instruction to testers. Weak: "our testers are careful." Care is not a control.
- What are your stop conditions, and who can invoke them? Good: named triggers such as evidence of prior compromise, service degradation or reaching sensitive data, with immediate escalation and either side able to call a halt. Weak: no defined stop conditions.
- If you find evidence of an existing compromise, what happens? Good: stop testing, preserve evidence, notify your named escalation contact immediately, and do not continue until you decide. Weak: "we would note it in the report."
- Do you test denial of service, social engineering or physical access? Good: clearly out of scope unless separately commissioned and separately authorised, with an explanation of why they need their own agreement. Weak: a casual yes without mentioning additional authorisation.
People and qualifications
- Who specifically will do this work, and can I meet them? Good: named individuals, their experience, and a willingness to put them on a call before contract. Weak: "one of our senior consultants" with no name until after signature.
- Is any part of this engagement subcontracted, and to whom? Good: a direct answer either way, and if yes, who they are, where they are and what they can access. Weak: hesitation, or an answer that only covers "delivery" while leaving triage or reporting unexplained.
- What certifications do your testers hold, and what do you do beyond certifications? Good: relevant credentials plus something substantive such as internal research, published advisories or a training programme. Weak: a list of acronyms with no supporting practice.
- What background checks and security training apply to people who can see our findings? Good: a defined pre-employment check, periodic renewal and role-based access so not everyone can see everything. Weak: an answer about testers only, ignoring support and engineering staff.
Tooling and automation
- What proportion of this engagement is automated, and what is manual? Good: a straight answer with the split explained by activity. Both a heavily automated platform and a heavily manual consultancy can answer this well. Weak: evasion, or a claim that everything is manual while the price says otherwise.
- How are findings validated before they reach us? Good: describes confirmation of exploitability, evidence capture and a false-positive reduction process, with an honest figure for what proportion is manually checked. Weak: "our scanner is very accurate."
- What does exploitation mean in your process, and what will you not do? Good: proof of impact with the minimum action necessary, no destructive operations, no data exfiltration beyond an evidence sample, all gated by the authorisation document. Weak: either no exploitation at all while selling a penetration test, or exploitation with no stated limits.
- How do you score severity? Good: CVSS as a baseline, an explicit statement of which version, an adjustment for business context, and real-world exploitation signals such as the known-exploited vulnerability catalogue from CISA. Weak: a proprietary score with no published basis and no mapping to anything you already use.
- If you use AI anywhere in this process, where does our data go? Good: a clear statement of where models run, whether scan data leaves their infrastructure, and whether any of it is used for training. Weak: an enthusiastic answer about AI capability that does not mention data handling at all.
Reporting and remediation
- Can I see a redacted sample report before we sign? Good: yes, promptly, and preferably one resembling your scope. Weak: reluctance, or a two-page marketing sample. This is the single most informative artefact in the whole evaluation.
- What evidence do you supply per finding? Good: request and response detail, screenshots or captured output, the exact affected asset, and reproduction steps an engineer can follow alone. Weak: a severity rating and a generic description of the vulnerability class.
- How specific is your remediation guidance? Good: instructions tied to your actual technology and configuration. Weak: "apply vendor patches" or "validate all input," which tells you the finding was detected rather than investigated.
- What happens when we dispute a finding? Good: a defined process with a named contact, a re-examination and a documented outcome either way. Weak: "we would look at it," or a report that has no mechanism for correction.
- What does the retest include, and how long do we have? Good: clear terms on whether verification is included, how many passes, whether the report is reissued and how closure is evidenced for an auditor. Weak: vagueness, which reliably becomes an invoice.
Data handling and residency
- Where will our findings, evidence and reports be stored? Good: a named country and a named hosting arrangement, offered before you have to ask twice. Weak: "in the cloud," or an answer about their office rather than their infrastructure.
- Who can access our data, including any offshore teams? Good: named locations, role-based access, multi-factor authentication and audit logging, with support and engineering access disclosed rather than glossed over. Weak: "only your delivery team," which is almost never literally true for a platform.
- How long do you retain our data, and how is deletion evidenced? Good: stated retention periods, an automated deletion process, and a route to request earlier removal or a legal hold. Weak: "as long as necessary."
- What certification or independent assurance does your own organisation hold? Good: something verifiable such as ISO 27001 with the scope statement, or an attestation report, plus a clear statement of what it does not cover. Weak: a trust badge on the website with nothing behind it, or an alignment claim presented as certification.
- What will you tell us, and how fast, if you are breached? Good: a contractual notification commitment with a stated period and named contact. Weak: no commitment, or one that only appears in a document nobody has signed.
Commercial and contractual
- What pricing model is this quote, and what would three years cost? Good: names the model, states what is chargeable, and models growth honestly. Weak: a single number with no model behind it and no view of year two.
- What is not included? Good: an immediate, specific list, typically covering additional roles, assets added later, out-of-hours work and repeat retests. Weak: "everything is included," which never survives contact with an actual engagement.
- Which contract documents can you sign, and how long does that take? Good: master services agreement, data processing agreement, non-disclosure and an authorisation-to-test record, with realistic turnaround. Weak: surprise at being asked for a data processing agreement.
- Can we speak to a reference in our sector? Good: yes, arranged within a reasonable period, ideally with someone facing your regulatory context. Weak: case studies offered instead of a conversation.
The four answers that should end the conversation
- A guarantee. Nobody can guarantee they will find everything, that you will pass an audit, or that you will be secure afterwards. A provider offering any of those is either inexperienced or hoping you are.
- Certification claims that will not survive a follow-up question. Ask for the scope statement of any certification named. Alignment with a framework is legitimate and useful; presenting it as certification is not.
- Refusal to show a sample report. Confidentiality is a reason to redact, not a reason to withhold. Every credible provider has a sanitised sample ready to send.
- No answer on data residency. They are about to hold a catalogue of your weaknesses. If they cannot say where it lives and who can read it, the rest of the evaluation does not matter.
How PentestOps answers these
Applying our own list to ourselves, and being specific rather than flattering.
Methodology and limits: a seven-phase process aligned to PTES, the OWASP Web Security Testing Guide v4.2, NIST SP 800-115, CREST testing guidance and CIS Benchmarks, covering external, internal, web application, API, cloud and identity. What is out of reach: deep multi-step business logic specific to your product, and judgement calls that need a human who understands your domain. Automation is a coverage layer, not a replacement for a skilled tester. See methodology.
Authorisation and stop conditions: every scan and exploitation action is gated by per-tenant rules of engagement, with scope enforcement that automatically stops activity outside authorised assets, and a full evidence trail. Assets are verified before they can be tested. See rules of engagement.
Tooling, AI and severity: findings are validated with safe automated exploitation that proves impact rather than asserting it, scored with CVSS v3.1 and prioritised using CISA KEV. Our AI runs self-hosted by default on infrastructure Extranet Systems operates, so scan data is not sent to third-party model providers; external providers can be enabled per tenant if a customer wants them. See AI in penetration testing.
Reporting and retest: evidence per finding, remediation guidance tied to the affected asset, mapping to 8 compliance reporting frameworks (OWASP Top 10, PCI DSS v4.0, NIST 800-53, SOC 2, HIPAA, GDPR, ISO 27001, SMB1001) plus CIS Benchmarks for AWS, Azure, GCP and Kubernetes, and multiple report formats including PDF, CSV and JSON/API. Because testing runs on a schedule, a fixed finding is re-verified on the next run rather than through a separately purchased retest.
Data handling: Extranet Systems Pty Ltd is ISO/IEC 27001:2022 certified, independently audited by Atom Assurances, and the certificate is available on request. The platform is built to SOC 2-aligned controls and a SOC 2 attestation is on our roadmap; we do not describe that as certification. The platform is hosted in Australia on infrastructure operated by Extranet Systems and customer data is stored in Australia. Authorised personnel in our Bahrain and Egypt offices may access customer data where necessary to operate, support and secure the platform. Retention runs from 1 year to 7 years by plan, with 365-day audit logs. See Trust Centre and DPA.
Commercials: asset-wise pricing, published rather than quoted, with scans against an asset unlimited within fair use. All paid plans start with a 7-day free trial. A card is required to start your trial and is only charged after the trial ends, unless you cancel first. See pricing.