From fuzzers to AI hackers: What seven years of application security research solved, and what remains unsolved
Giorgio Roffo, Alessio Dalla Piazza
Table of contents
- The trajectory looks finished. It isn’t.
- Seven years in five leaps
- The half the field solved and the half it left unsolved
- Coverage was never the point
- What counts as a verified exploit
- A frontier model is not a pentest platform
- Where Equixly is today
- Building for the missing half
- Measuring the missing half
- Conclusion: Closing the missing half
The field has learned to generate, sequence, and even
reason about correct requests. It still hasn’t learned
to prove an exploit, run safely forever, or check its
own fixes.
Abstract
Between 2019 and 2026, automated REST API testing went from “emit a valid request from a schema” to “coordinate several LLM-driven agents.” That progress is real, but it is lopsided. Research has largely solved the left half of the offensive security life cycle (discover, understand, and generate) and barely touched the right half (verify a reproducible exploit, reason across multiple identities, run continuously and safely, and retest after a fix). The frontier is no longer a smarter fuzzer or a bigger model. It is a reproducible, governed, continuously operating system that produces verified security evidence. That is a machine learning and systems problem, which is exactly the direction Equixly’s research pursues, while the literature is still catching up.
The trajectory looks finished. It isn’t.
From the abstracts alone, application security testing looks like a solved problem waiting for compute. In 2019, the state of the art was generating a syntactically valid request from an OpenAPI document. By 2026, there are systems that coordinate multiple agents, train reinforcement-learning policies, and rank each other in tool competitions.
The results sections tell a different story. Almost every system stops at the same place. It can find anomalies (a 500, a schema deviation, or a suspicious response). However, it cannot, on its own, turn a suspicion into a verified, reproducible exploit, operate safely and continuously against a system that changes under it, or retest once a fix ships.
Applications and their APIs are now the main enterprise attack surface, and the OWASP API Security Top 10 (2023) is dominated by authorization logic flaws, such as Broken Object Level Authorization (BOLA), Broken Function Level Authorization (BFLA), and mass assignment (now part of Broken Object Property Level Authorization, or BOPLA).
These are not memory-safety bugs. A successful attack usually returns a perfectly well-formed 200 OK. You cannot detect them by watching for a crash.
You detect a BOLA by comparing responses across two identities. You detect excessive data exposure by differencing responses under field suppression. And you detect an access-policy violation by reasoning about what a defined role should be allowed to see.
Seven years in five leaps
The research falls into five overlapping eras, each building on the last.
Figure 1. Five eras of automated REST API security testing, 2019–2026. Each band rises with increasing autonomy, security evidence, and operational readiness.
Era I: Schema-driven generation (2019–2022)
The founding idea: an OpenAPI document is a machine-readable contract you can derive tests from.
QuickREST (ICST 2020) generated property-based tests with shrinking. RESTTestGen (ICST 2020) built an operation-dependency graph. RESTest (ISSTA 2021) added a constraint solver for inter-parameter dependencies. Schemathesis (ICSE 2022) derived structure- and semantics-aware fuzzers from the schema.
Two limitations carried forward:
- Oracles were weak (status codes, schema conformance).
- RESTTestGen excluded from its evaluation any API that required authentication; even logging in was, at the time, out of scope.
Era II: Stateful and dependency-aware fuzzing (2019–2025)
Real APIs are stateful: you must POST a resource before you can GET or DELETE it.
RESTler (ICSE 2019), “the first stateful REST API fuzzer,” inferred producer–consumer dependencies and pruned invalid sequences from response feedback (28 new GitLab bugs, plus production Azure findings). MOREST (ICSE 2022), RestCT (ICSE 2022), and foREST (ISSRE 2023) refined the dependency model. The community’s own benchmarks (Open Problems in Fuzzing, TOSEM 2023; No Time to Rest Yet, ISSTA 2022) confirmed that white-box EvoMaster (Automated Software Engineering journal, 2025) led on coverage.
What this era did not deliver: authentication was, at best, a single static token. No system reasoned about multiple identities or privilege boundaries, which is the exact ingredient that authorization bugs require.
Era III: Vulnerability-oriented testing (2020–2026)
Here, the question shifted from “did it crash?” to “is this specific vulnerability present?”
The lineage began inside Era II. Atlidakis et al.’s security checkers (ICST 2020) reissued requests under a second user’s token to test cross-tenant isolation, the earliest first-class multi-identity oracle.
Then the field diversified by class:
- NAUTILUS (USENIX Security 2023), with 23 zero-days including a Confluence RCE
- MINER (USENIX Security 2023)
- Mass-assignment testing (Corradini et al., ICSE 2023)
- EDEFuzz (ICSE 2024), the first oracle for excessive data exposure
- VOAPI2 (USENIX Security 2024), with SSRF/injection payloads and 23 CVEs
The most complete authorization work to date, Arcuri et al.’s access-policy oracles (ISSRE 2025), explicitly requires multiple authenticated roles to catch BOLA/BFLA. However, it is still evaluated on vulnerable-by-design benchmarks, not production, and neither runs continuously nor retests.
When Corradini et al. applied their mass assignment testing to ten production Google APIs, they deliberately stopped at static candidate identification, “to avoid the risk of mounting a successful injection attack on a service running in production.” Safe, non-destructive operation against production systems remains an open research problem.
Era IV: Learning- and LLM-guided testing (2024–2025)
The target, this time, was input quality.
RESTGPT (ICSE 2024) prompts an LLM to mine rules from the natural language parts of a spec (valid inputs for 73% of parameters vs. 17% for a prior baseline). APIRL (AAAI 2025) trains a deep-RL fuzzer with genuine train/test API separation and repeated runs (the most ML-rigorous system in the corpus). Yet, its findings are still server errors, not verified exploits.
LlamaRestTest (FSE 2025) took a different route: it fine-tuned and quantized two small models and beat the large-prompted-LLM baselines. A small, specialized, quantized model can outperform a big, general, prompted one on a domain task.
Era V: Multi-agent and agentic systems (2025–2026)
The newest work coordinates several components.
AutoRestTest (ICSE 2025) combines multi-agent RL, a semantic dependency graph, and LLM-generated values, and topped a 2026 tool competition. Test-amplification pipelines (Nooyens et al., ICTSS 2025) and metamorphic-testing agents like ARMeta (COMPSAC 2026) add more agents.
But “multi-agent” here means agents cooperating on one testing objective (coverage, amplification, oracles), not agents dividing the offensive life cycle into discover/plan/exploit/verify roles.
None of these systems produces a deterministically verified exploit, reasons across identities at scale, runs in CI/CD, or retests after remediation.
Figure 2. The 22 representative systems placed on the 2019–2026 timeline by era. Each system is marked with the ingredient it added: schema-driven generation, stateful sequencing, a vulnerability-specific oracle, learned or LLM-guided inputs, or multi-agent coordination.
The half the field solved and the half it left unsolved
Scoring the representative systems against the same twelve capabilities makes the imbalance visible.
Figure 3. 22 systems scored on 12 capabilities. The dashed line splits the broadly solved capabilities (generation, state, and learning) from the operational ones (verification, continuous operation, and remediation). The right half is nearly empty for every existing system.
The left of that grid is solved: request generation, producer–consumer dependencies, and stateful sequencing (RESTler, MOREST, RestCT, foREST, EvoMaster). Vulnerability-oriented work produced real zero-days (NAUTILUS, VOAPI2) and gave logic-vulnerability classes their first oracles (the cloud-service checkers, mass assignment, EDEFuzz, and the access-policy oracles).
The right of the grid is where the work stops:
- Verified exploitation. A handful of systems produced vendor-confirmed CVEs (RESTler, the cloud-service checkers, NAUTILUS, and VOAPI2), but always through manual triage, never as a native, replayable proof artifact inside a closed loop. No system in the corpus outputs a deterministic, re-runnable exploit as its unit of evidence.
- Multiple identities and business logic. Solved for producer–consumer state, shallow for single-token auth, and rare for the multi-identity reasoning that BOLA/BFLA demand (the cloud-service checkers, APIRL, and Arcuri et al.’s access-policy oracles). Genuine multi-step business workflows are barely represented.
- Repeated-run reliability. Quantified in only a few papers (Schemathesis, mass assignment, and APIRL). A study that ran 400 autonomous pentests against one fixed target found success rates swinging from 25% to 85% depending on the model (Erdem, 400-Run Empirical Study, arXiv 2026), a warning against trusting any single run.
- Continuous, safe operation and remediation retesting. Continuous operation appears in essentially one system (RESTest). Safe production exploitation is so unsolved that researchers declined to try it (the mass-assignment work), and retesting after a fix is absent from every system surveyed here.
The broader agentic pentest literature reaches the same verdict from the outside: surveys flag weak multi-stage performance and unreliable evaluation (He et al., A Survey of LLM-Driven Penetration Testing, arXiv 2026). Controlled studies show that planning and state-management failures persist regardless of tooling (Deng et al., What Makes a Good LLM Agent for Real-world Penetration Testing?, arXiv 2026).
Equixly’s research is directed at this right half. No published method addresses the gray cells (deterministic exploit verification, multi-identity and business-logic reasoning, continuous CI/CD operation, and remediation retesting) end to end, and they are the ones that matter most in practice.
The rest of this piece explains what closing them requires.
Coverage was never the point
The right half of that grid is empty because the field has been optimizing a different loop from the one enterprises run.
Figure 4. The academic loop (left) maximizes coverage and unique server errors. The enterprise loop (right) is longer and closes on itself: discover, map dependencies, establish authenticated state, form an attack hypothesis, exploit safely, reproduce deterministically, report, and retest after deployment.
The academic loop (generate request → observe response → increase coverage) is an effective engine for finding robustness defects, and the field has tuned it well (Open Problems in Fuzzing, TOSEM 2023; EvoMaster, Automated Software Engineering journal, 2025). But coverage and unique 500s are only weakly connected to exploitable business risk.
A tool can maximize coverage and find nothing exploitable. It can emit thousands of “potential” issues (EDEFuzz, ICSE 2024) that a human must confirm; it can win a competition on error count (AutoRestTest, ICSE 2025) and still say nothing about whether one reported issue is a reproducible exploit an enterprise must fix.
In practice, that gap is a triage cost: a scanner that emits thousands of unconfirmed “potential” findings burns analyst hours on triage. Conversely, a system that returns a replayable exploit with proof says exactly what to fix, and can confirm later that the fix held.
The enterprise loop insists on authenticated, multi-identity state. It is driven by attack hypotheses rather than coverage, demands deterministic reproduction as evidence, and closes on itself.
After a fix ships, the system has to decide whether the vulnerability is still exploitable, distinguishing partial remediation from complete remediation. Running that loop continuously and safely against a moving target is the real job.
What counts as a verified exploit
The component the literature lacks, and the one everything else depends on, is a deterministic verifier. It turns a model’s hypothesis (“this endpoint appears vulnerable”) into replayable proof:
exact request sequence + authentication context + expected insecure behavior + a control request + observed evidence + a replayable artifact
The verifier (not the language model) decides whether a reportable vulnerability exists. That distinction is what converts a suspected finding (the norm across the corpus) into verified evidence. It is also what makes false-positive control tractable and what makes remediation retesting well-defined: after a fix, you replay the same artifact. The exploit either still succeeds (partial fix) or no longer succeeds (complete fix).
A frontier model is not a pentest platform
A capable enough model does not close the gap on its own. The reliability evidence (Erdem, arXiv 2026; Deng et al., arXiv 2026) shows that scaling the model neither guarantees reproducibility nor removes planning failures. The more useful framing is architectural.
Figure 5. A general-purpose model supplies exactly one layer: reasoning. Planning and memory, security skills, typed tools, scope and safety controls, a session and multi-identity manager, a deterministic verifier, and the reporting/retest loop are engineered, trained, and governed.
A frontier model can increasingly find an exploit. What matters is what surrounds the model so that its finding becomes verified, reproducible, safe, and continuously recheckable evidence.
Every layer above the model maps to a gap from the grid:
- Typed tools and a session/identity manager cover multi-identity testing.
- Scope and safety controls address the destructive-action problem (mass assignment, ICSE 2023).
- The deterministic verifier addresses the evidence problem.
- The reporting-and-retest loop closes the enterprise loop.
Where Equixly is today
Equixly brings agentic security testing into the software delivery life cycle. Its agents explore applications and APIs, chain interactions across endpoints, and adapt their attacks to observed behavior.
This approach addresses vulnerabilities whose impact depends on application context: which user is acting, which resources they can access, and how several individually legitimate operations can combine into an exploitable workflow.
For security and engineering teams, the practical value is a clearer path from discovery to remediation. Equixly connects continuous testing with validated findings, helping teams understand how a weakness can be exploited, prioritize what to fix, and validate remediation as changes are deployed. Integrating this process into development workflows makes security testing part of how an application evolves.
Equixly AI Research builds on this foundation by addressing the problems that determine how effectively the platform can operate: reasoning through longer attack sequences, testing across identities and permissions, strengthening exploit verification, and developing specialized models that make frequent testing more efficient.
These priorities connect research progress to practical outcomes: more reliable findings, less manual investigation, and faster feedback for developers. The following sections explain the methods and evaluation principles guiding that work.
Building for the missing half
The research gaps identified earlier define a research agenda, and the one Equixly’s ML research is directed at: a deterministic verifier, multi-identity and business logic reasoning, continuous and safe operation, and remediation retesting. These are the capabilities the published state of the art leaves in the gray cells.
What follows is what closing each gap completely requires and the direction Equixly takes toward it, stated as design and training principles rather than an experimental report. This article presents no benchmarks and claims none. Instead, it sets out how the problem should be framed and the approach Equixly pursues.
Rigorous, governed data-to-model life cycle
Since the differentiator is a trained system rather than a prompt, the standard that matters is the machine learning life cycle, the same one any credible ML result is held to.
Figure 6. Data → curation → training → validation → deployment, with a one-way governance boundary. Explicit train/validation/test separation, negative controls, ablations, and evaluation on held-out realistic applications, and customer execution that produces evidence, not training data.
Several systems in the literature train and evaluate on the same target (FuzzTheREST, DCAI 2024) or leave the held-out split unclear (LlamaRestTest, FSE 2025), which makes generalization claims fragile.
APIRL (AAAI 2025) shows the standard can be met even here. The governance boundary (customer execution never flowing back into training) is a design principle. Specific retention guarantees belong in a public technical statement, not a blog.
Reinforcement learning aimed at verified exploitation
The bottleneck here is not model size but what the agent is optimized for.
A policy trained to maximize coverage or server errors becomes very good at poking an API; the objective that actually matters is a verified exploit.
The field is already moving in this direction, from RL fuzzers (APIRL, FuzzTheREST) to what the pentest survey calls Reinforcement Learning with Verifiable Rewards (He et al., arXiv 2026). This is the direction Equixly’s research pursues: framing the training signal around verified exploitation rather than coverage.
Figure 7. An agent policy acts on a purpose-built cyber range (seeded vulnerabilities, authentication, multiple identities, and business workflows) trained with policy-gradient methods (PPO, Schulman et al. 2017; GRPO, Shao et al. 2024). The reward is a verified exploit or completed attack chain. Destructive or out-of-scope actions are penalized.
Two design choices define the approach:
- The environment is a cyber range that instantiates exactly the capabilities the corpus lacks (multiple identities, business logic, and database state), so progress there requires reasoning about them.
- The reward is verifiable by design: an exploit should count only if the deterministic verifier can replay it, which makes the evidence oracle the training signal itself.
Penalizing destructive and out-of-scope actions is how safety lives in the objective rather than being bolted on afterward.
Efficient models for deeper security testing
The cost of an agentic scan accumulates across the entire investigation. Every unusable request, repeated attempt, and incorrect interpretation can consume resources without advancing the test. For Equixly, improving model efficiency means reducing the work needed to reach a reliable security conclusion.
Specialization offers a practical route. Tasks such as identifying rejected parameters or classifying response outcomes have defined inputs and measurable answers, making them candidates for compact models.
Planning an attack across several endpoints and identities requires broader reasoning. Evaluating these tasks separately helps determine where smaller models are sufficient and where additional reasoning improves results.
LlamaRestTest provides evidence for this approach: its specialized models improved valid-input generation and parameter-dependency detection in the REST APIs evaluated (LlamaRestTest, FSE 2025).
Distillation and quantization address different parts of the efficiency problem:
- Distillation trains a student model using supervision from a teacher, potentially transferring useful behavior into a smaller model.
- Quantization reduces numerical precision to lower memory requirements and, with suitable hardware support, to accelerate inference.
Quantization-aware training lets the model adapt to those numerical constraints during training. These techniques provide concrete research directions for deploying specialized security models within a manageable computing budget [Hinton et al., 2015].
Figure 8. Distillation transfers learned behavior into a compact model; quantization reduces its numerical precision. The resulting model must be evaluated for security effectiveness, memory use, and inference speed.
The decisive evaluation is what happens to the security test after these changes. On previously unseen applications, does the system still detect the same vulnerabilities? Does it introduce false positives, require more retries, or abandon difficult paths? Measuring these outcomes alongside latency and memory establishes whether compression improves the platform overall.
A manageable model footprint also broadens deployment options. Where customer requirements call for local inference, self-hosting gives control over where the model runs and when it is updated. Versioned deployments allow each proposed update to be evaluated against a fixed security benchmark before rollout.
It also addresses the concern raised by the reliability literature (Erdem, arXiv 2026; Deng et al., arXiv 2026): a system whose intelligence layer can be pinned and self-hosted is far easier to evaluate with the repeated-run rigor those studies demand.
Together, these four (a governed life cycle, a policy rewarded for verified exploitation, a self-hostable model, and a deterministic verifier) are what an enterprise security program needs, and what a thin wrapper around a third-party model cannot replicate: findings that arrive with proof, testing that runs daily rather than quarterly, deployment that respects data residency, and an audit trail for every action.
That is why Equixly treats the gray half of the map as its core problem rather than a feature.
Measuring the missing half
Measuring any of this takes more than coverage and error counts, the metrics that decide tool competitions such as the SBFT 2026 REST League that AutoRestTest topped. Those are necessary but not sufficient.
The metric families that track the operational half of the life cycle are the ones the field mostly doesn’t report yet:
- Exploitation: verified-exploit success rate, attack-chain completion, time and requests to first verified exploit
- Reliability: success probability and variance across N runs; premature-termination, hallucinated-tool, and context-exhaustion rates (Erdem, arXiv 2026; Deng et al., arXiv 2026)
- Efficiency: cost, tokens, and requests per verified vulnerability
- Safety: destructive-action rate, scope and rate-limit violations, and audit-log completeness
- Remediation: successful-retest rate, regression detection, and partial-vs-complete remediation
Adopting those, especially the verified-exploit and repeated-run families, would let the field measure the half of the job it has so far left unmeasured.
Conclusion: Closing the missing half
Automated API security testing went from generating a valid request to coordinating reinforcement-learning agents in seven years. It solved generation, dependency discovery, and statefulness. It demonstrated vulnerability-specific oracles. And it began to reason and act in concert.
Measured against the complete offensive-security life cycle (discover, understand, plan, exploit, verify deterministically, report, and retest after remediation, as well as run continuously and safely), it stops, almost universally, at the halfway point.
So the frontier isn’t a smarter fuzzer or a bigger model. It’s a reproducible, governed, continuously-operating system that produces verified security evidence. It comprises purpose-built environments, reinforcement learning rewarded for verified exploitation, efficient self-hostable models, and a deterministic verifier that makes evidence replayable.
That is the missing half, and the one Equixly’s research is built to close: turning a promising capability into evidence a security team can trust.
References
[1] OWASP Foundation, “OWASP API Security Top 10 (2023),” 2023. [Online]. Available: https://owasp.org/API-Security/
[2] S. Karlsson, A. Čaušević, and D. Sundmark, “QuickREST: Property-based test generation of OpenAPI-described RESTful APIs,” in Proc. IEEE Int. Conf. Softw. Testing, Validation and Verification (ICST), 2020, pp. 131–141.
[3] E. Viglianisi, M. Dallago, and M. Ceccato, “RESTTestGen: Automated black-box testing of RESTful APIs,” in Proc. IEEE Int. Conf. Softw. Testing, Validation and Verification (ICST), 2020, pp. 142–152.
[4] A. Martin-Lopez, S. Segura, and A. Ruiz-Cortés, “RESTest: Automated black-box testing of RESTful web APIs,” in Proc. ACM SIGSOFT Int. Symp. Softw. Testing and Analysis (ISSTA), 2021, pp. 682–685.
[5] Z. Hatfield-Dodds and D. Dygalo, “Deriving semantics-aware fuzzers from web API schemas,” in Proc. IEEE/ACM Int. Conf. Softw. Eng. Companion (ICSE-Companion), 2022, pp. 345–346, arXiv:2112.10328.
[6] V. Atlidakis, P. Godefroid, and M. Polishchuk, “RESTler: Stateful REST API fuzzing,” in Proc. IEEE/ACM Int. Conf. Softw. Eng. (ICSE), 2019, pp. 748–758.
[7] Y. Liu et al., “MOREST: Model-based RESTful API testing with execution feedback,” in Proc. IEEE/ACM Int. Conf. Softw. Eng. (ICSE), 2022, pp. 1406–1417.
[8] H. Wu, L. Xu, X. Niu, and C. Nie, “Combinatorial testing of RESTful APIs,” in Proc. IEEE/ACM Int. Conf. Softw. Eng. (ICSE), 2022, pp. 426–437.
[9] J. Lin et al., “foREST: A tree-based black-box fuzzing approach for RESTful APIs,” in Proc. IEEE Int. Symp. Softw. Reliability Eng. (ISSRE), 2023, pp. 695–705.
[10] M. Zhang and A. Arcuri, “Open problems in fuzzing RESTful APIs: A comparison of tools,” ACM Trans. Softw. Eng. Methodol., vol. 32, no. 6, 2023.
[11] M. Kim, Q. Xin, S. Sinha, and A. Orso, “Automated test generation for REST APIs: No time to rest yet,” in Proc. ACM SIGSOFT Int. Symp. Softw. Testing and Analysis (ISSTA), 2022, pp. 289–301.
[12] A. Arcuri et al., “EvoMaster: Black and white box search-based fuzzing for REST, GraphQL and RPC APIs,” Automated Softw. Eng., vol. 32, no. 4, 2025.
[13] V. Atlidakis, P. Godefroid, and M. Polishchuk, “Checking security properties of cloud service REST APIs,” in Proc. IEEE Int. Conf. Softw. Testing, Validation and Verification (ICST), 2020, pp. 387–397.
[14] G. Deng et al., “NAUTILUS: Automated RESTful API vulnerability detection,” in Proc. USENIX Security Symp., 2023, pp. 5593–5609.
[15] C. Lyu et al., “MINER: A hybrid data-driven approach for REST API fuzzing,” in Proc. USENIX Security Symp., 2023, pp. 4517–4534, arXiv:2303.02545.
[16] D. Corradini, M. Pasqua, and M. Ceccato, “Automated black-box testing of mass assignment vulnerabilities in RESTful APIs,” in Proc. IEEE/ACM Int. Conf. Softw. Eng. (ICSE), 2023, pp. 2553–2564, arXiv:2301.01261.
[17] L. Pan, S. Cohney, T. Murray, and V.-T. Pham, “EDEFuzz: A web API fuzzer for excessive data exposures,” in Proc. IEEE/ACM Int. Conf. Softw. Eng. (ICSE), 2024, arXiv:2301.09258.
[18] W. Du, J. Li, Y. Wang, L. Chen et al., “Vulnerability-oriented testing for RESTful APIs (VOAPI2),” in Proc. USENIX Security Symp., 2024, pp. 739–755.
[19] A. Arcuri, O. Sahin, and M. Zhang, “Fuzzing for detecting access policy violations in REST APIs,” in Proc. IEEE Int. Symp. Softw. Reliability Eng. (ISSRE), 2025, pp. 130–141; extended in O. Sahin, M. Zhang, and A. Arcuri, “Enhancing REST API fuzzing with access policy violation checks and injection attacks,” arXiv:2604.00702, 2026.
[20] M. Kim, T. Stennett, D. Shah, S. Sinha, and A. Orso, “Leveraging large language models to improve REST API testing (RESTGPT),” in Proc. IEEE/ACM Int. Conf. Softw. Eng.: New Ideas and Emerging Results (ICSE-NIER), 2024, arXiv:2312.00894.
[21] M. Foley and S. Maffeis, “APIRL: Deep reinforcement learning for REST API fuzzing,” in Proc. AAAI Conf. Artif. Intell., vol. 39, no. 1, 2025, pp. 191–199, arXiv:2412.15991.
[22] M. Kim, S. Sinha, and A. Orso, “LlamaRestTest: Effective REST API testing with small language models,” Proc. ACM Softw. Eng., vol. 2, no. FSE, art. FSE022, 2025, arXiv:2501.08598.
[23] M. Kim, T. Stennett, S. Sinha, and A. Orso, “A multi-agent approach for REST API testing with semantic graphs and LLM-driven inputs (AutoRestTest),” in Proc. IEEE/ACM Int. Conf. Softw. Eng. (ICSE), 2025, arXiv:2411.07098; see also T. Stennett et al., “AutoRestTest at the SBFT 2026 tool competition,” arXiv:2607.01063, 2026.
[24] R. Nooyens, T. Bardakci, M. Beyazıt, and S. Demeyer, “Test amplification for REST APIs via single and multi-agent LLM systems,” in Proc. IFIP Int. Conf. Testing Softw. and Systems (ICTSS), 2025, arXiv:2504.08113.
[25] S. Khan, A. Mughees, G. Sudheerbabu, T. Ahmad, and D. Truscan, “Multi-agent LLM-based metamorphic testing for REST APIs (ARMeta),” in Proc. IEEE Annu. Computers, Software, and Applications Conf. (COMPSAC), 2026, arXiv:2605.28321.
[26] G. T. Erdem, “How reliable are AI attackers against a fixed vulnerable target? A 400-run empirical study of LLM penetration testing consistency,” arXiv:2605.30096, 2026.
[27] Z. He et al., “A survey of LLM-driven penetration testing: Taxonomy, co-evolution, and open challenges,” arXiv:2607.02605, 2026.
[28] G. Deng et al., “What makes a good LLM agent for real-world penetration testing?,” arXiv:2602.17622, 2026.
[29] T. Dias, E. Maia, and I. Praça, “FuzzTheREST: An intelligent automated black-box RESTful API fuzzer,” in Proc. Int. Conf. Distributed Computing and Artificial Intelligence (DCAI), 2024, arXiv:2407.14361.
[30] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv:1707.06347, 2017.
[31] Z. Shao et al., “DeepSeekMath: Pushing the limits of mathematical reasoning in open language models (GRPO),” arXiv:2402.03300, 2024.
[32] G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv:1503.02531, 2015.
Giorgio Roffo
Head of AI
Giorgio is an AI leader with a Ph.D. in Computer Science, focused on machine learning and pattern recognition. His expertise spans computer vision, scalable AI systems, and applied machine learning. He has worked across industry and academia, translating advanced research into reliable production technology. His work includes medical AI, computer vision, and security-focused intelligent systems, with publications in top-tier international venues and multiple research awards.
Alessio Dalla Piazza
CTO & FOUNDER
Former Founder & CTO of CYS4, he embarked on active digital surveillance work in 2014, collaborating with global and local law enforcement to combat terrorism and organized crime. He designed and utilized advanced eavesdropping technologies, identifying Zero-days in products like Skype, VMware, Safari, Docker, and IBM WebSphere. In June 2016, he transitioned to a research role at an international firm, where he crafted tools for automated offensive security and vulnerability detection. He discovered multiple vulnerabilities that, if exploited, would grant complete control. His expertise served the banking, insurance, and industrial sectors through Red Team operations, Incident Management, and Advanced Training, enhancing client security.