AI is changing the economics of cyberattacks. What earlier required a skilled researcher spending hours on reconnaissance, testing and adapting an attack can increasingly be done by AI agents at much greater speed and scale. But most enterprise penetration testing programs have not changed at the same pace. We still test a limited set of applications, often once or twice a year, using a combination of human testers and traditional automated tools. I believe this model is going to change significantly. If attackers can test more assets, more frequently, and connect weaknesses across applications, APIs and infrastructure, defenders will have to rethink how penetration testing itself is designed.
The old penetration testing model is starting to break
The first issue is scope. Most organizations focus penetration testing on the crown jewels. That is understandable when capacity is limited. But an attacker does not have to start there. A forgotten application, exposed API, staging environment or weak credential can provide the first foothold, and from there the attacker can move laterally. The weakest asset can therefore become the path to the most valuable asset. Testing a critical application in isolation is not enough. We have to ask whether there is any practical path across the broader environment that can eventually reach something critical.
Also read: India Sees 3,237 Cyberattacks Weekly per Organization: Check Point Research
The second issue is frequency. A penetration test is a snapshot, while the attack surface keeps changing. New applications go live, APIs change, credentials leak and new vulnerabilities appear. Testing once or twice a year increasingly leaves a large gap between what was tested and what exists today. The third issue is scale. Human pentesters are very good at depth and understand context, business logic and unusual application behavior. But a human team cannot deeply test hundreds or thousands of assets often enough. The mathematics simply does not scale.
Automation helped, but solved only part of the problem
Automation did solve part of the scalability problem. Vulnerability scanners and DAST tools made it possible to test large numbers of applications much faster than humans could. ASM improved visibility into what was exposed, while BAS helped teams simulate attacker techniques and validate whether controls detected or blocked them.
But traditional web application testing automation had an important limitation. Many classes of vulnerabilities were difficult to automate because they required context and reasoning. Business-logic vulnerabilities, IDOR, BOLA, authorization issues and multi-step application attacks often needed a human tester to understand how the application was supposed to behave and then work out how to break that logic.
This created a trade-off. Automated tools provided breadth and frequency, while human pentesters provided depth. AI agents are interesting because, for the first time, we are beginning to see these two worlds come together.
What reaching No. 1 on HackerOne taught us about AI pentesting
We wanted to test this hypothesis in an environment where AI agents had to compete with skilled human researchers. HackerOne gave us exactly that environment.
What was particularly interesting was not only that the agents found vulnerabilities. They found zero-day vulnerabilities in applications that had already gone through existing security testing, and in several cases before human researchers working on the same targets. Over one quarter, our agents reached No. 1, No. 2 and No. 3 positions across different HackerOne leaderboard categories, with every submission reviewed by a human before it went out.
The economics were equally important. The agents ran at roughly $5,000 per month in cloud and AI compute, comparable to a single manual penetration test and materially below the cost of adding enough human capacity to test a large attack surface frequently. This changes how much depth an enterprise can afford.
Building an optimal pentesting program: balancing depth, breadth, and frequency
I think the right way to design the next testing program is not to maximize any one of these dimensions. It is to optimize the balance between all three. We need to start by discovering the attack surface and classifying assets based on criticality, exposure and how frequently they change. Then decide both testing depth and cadence based on that risk. A critical internet-facing application may justify deep AI-led pentesting much more often, especially after a major release or material change. A lower-risk asset may go through lighter testing more frequently and deeper testing less often. A newly exposed API or critical vulnerability may trigger testing immediately, regardless of the normal schedule.
The objective is not maximum depth everywhere. It is the right depth, on the right assets, at the right frequency. This also makes the model practical as AI testing becomes increasingly consumption-driven, because deeper agentic testing has real compute and token cost. The testing program should behave more like a risk-based portfolio than a single calendar-driven exercise.
Historically, organizations have had to trade these dimensions against each other. Human testing gives depth, but limited breadth and frequency. Traditional automation gives breadth and frequency, but often less depth. AI creates the possibility of getting much closer to all three.
The future is AI + humans, with safety and governance built in
The HackerOne results do not mean the human pentester is finished. I believe the opposite. AI changes where human expertise creates the most value. AI agents are well suited for reconnaissance, enumeration, repetitive testing, payload experimentation, vulnerability validation and attack-path exploration at a scale humans cannot match. Humans should focus where judgment matters most: complex business logic, ambiguous authorization, understanding business impact, defining objectives and making high-impact decisions.
There is also a third piece that becomes increasingly important as offensive agents become more autonomous: safety and governance. These agents are taking real actions against real systems. Explicit scope controls, approval gates for high-impact actions, sandboxing, audit logs and a kill switch have to be part of the architecture, not an afterthought. In my view, the model and the agent are only part of the technology. The harness around the agent will become equally important.
So I do not see the future as AI replacing penetration testers. I see AI giving human experts the breadth and frequency they never had before, while humans continue to provide the depth, judgment and accountability that machines still struggle with.
The article has been written by Bikash Barai, Founder and CEO, FireCompass















