Knowledge documents: When your application rejects everything an agent tries
Alessio Dalla Piazza, Ion Farima, Nicolò Piccoli
Table of contents
Your testing agent hits an endpoint that registers lab samples. It tries sample_type: “test”. Rejected. It tries sample_type: “blood_sample”. Rejected. It tries sample_type: “A”. Rejected.
It will never succeed. The endpoint accepts exactly four values (BLOOD, TISSUE, CSF, and PLASMA), and those codes appear nowhere in the API specification. The agent cannot read them off the spec, and it cannot guess them, so every request it sends is rejected before the application logic ever runs.
A different agent, on a different API, has the opposite problem. Every value it sends is perfectly valid; it just calls POST /orders/{id}/pay before POST /orders/{id}/reserve, and the API rejects it anyway. Nothing in the request was wrong. A precondition had not happened yet, and nothing in the specification or the response body said so.
The instinct is to fix this with volume. That is what classic DAST does: spray and pray. Take a parameter, fire a payload list at it, and see what comes back.
Against a web form with a handful of loosely typed fields, that works well enough, because volume stands in for understanding. Against an API that only accepts values from a list it never publishes, it cannot work at all. The list was never in the payload file, and sending that file twice does not put it there.
Equixly’s testing agent does not spray; it adapts. It reads the response to every request, works out what the API objected to, and builds the next attempt differently. It does that in the way a pentester treats a 422 Unprocessable Entity, as a hint rather than a dead end: read which field the body names, change that one, send it again.
That covers far more ground than volume ever does. A malformed date, a missing required field, an identifier that had to come from an earlier call: all of those are recoverable, because the API’s own errors carry enough signal to work backward from.
What no amount of feedback produces is information that never appears in the exchange. BLOOD cannot be derived from a rejection that only says the specimen category is invalid. Adaptation needs something to adapt from, and when an API keeps its vocabulary to itself, there is nothing for the loop to converge on. The same is true of the payment that ran too early: the failure says the request was refused, not what should have run before it.
That is the core problem this article is about. An agent can only test what it can derive, and an API’s most important endpoints usually gate on things a specification never states, such as internal codes, compound identifiers, lookup tables, header secrets, and undocumented call order. A knowledge document is how you hand the agent what it is missing, whether that is a value or a sequence, so it can build requests the API accepts, in the order it accepts them.
The following sections explain:
- The five shapes that gap takes
- How the engine turns a document into something the testing agent uses at scan time
- How missing domain knowledge affects endpoint reachability and workflow execution
The endpoints an agent cannot reach
Every non-trivial API has endpoints that resist automated testing. Not because they are technically complex, but because they enforce internal rules that no external tool can derive on its own. They tend to fall into five shapes.
1. Hidden codes
The API requires a specific value from an internal list, but the documentation says merely “string.” Nothing tells the agent that collection_site must be a hospital code such as GH01-EMER, or that format only accepts HL7, FHIR, PDF, or RAW.
2. Compound format patterns
A field is not a free string; it is built by combining several codes with a separator. SMP-20240115-0042 is a sample ID: a prefix, a date, and a zero-padded sequence number. Without the pattern, an agent sends random strings and gets rejected every time.
3. Cross-field dependencies
The valid value of one field depends on another, with a lookup table between them. A spectrophotometer requires calibration standard NIST-SRM1921, and a centrifuge requires NIST-SRM2490. Send the wrong pair, and the API rejects it. The agent does not have the table.
4. Custom headers and access codes
Beyond standard authentication, some endpoints require extra headers with specific values, such as a facility code, a lab-section identifier, and an internal credential. Without them, even a correctly authenticated request is refused.
5. Execution order
Some operations must run in a specific sequence for reasons neither the specification nor the API’s own data flow exposes. An order can only be paid for once its stock has been reserved, even though the pay and reserve calls share no parameter that would let an agent infer the dependency. Both calls can be individually valid, and the later one still fails, because a precondition had not happened yet, rather than because an input was wrong.
For scale: in Equixly’s empirical benchmark (Dalla Piazza, 2025), automated tools produced a 93.5% failure rate across 86,310 requests, mostly because the inputs were values the API would never accept. Treat that number as an indication of the size of the problem, not as evidence that any one fix addresses it.
What a knowledge document contains
A knowledge document is plain documentation describing the rules an API enforces. You are not writing code or learning a specification format. You write things like an enum of accepted values:
Accepted specimen categories are: whole blood (BLOOD),
biopsy or surgical tissue (TISSUE), cerebrospinal fluid
(CSF), and separated plasma (PLASMA). No other specimen
descriptors are accepted.
A field structure:
The identifier follows the pattern SMP-{date}-{sequence}
where the date is eight digits (YYYYMMDD), and the
sequence is a zero-padded four-digit counter. For
example: SMP-20240115-0042.
A lookup table:
| Equipment Type | Required Calibration Standard |
|---|---|
| SPEC (Spectrophotometer) | NIST-SRM1921 |
| CENT (Centrifuge) | NIST-SRM2490 |
| PCR (Thermal cycler) | NIST-SRM2366 |
| MICRO (Microscope) | NIST-SRM2890 |
Or a required header:
Every request to a protected resource must include
the X-Facility-Code header. The current access code
is EQUIXLY-SECRET-42.
Or an execution order the API enforces but never states in its own responses:
Stock must be reserved before an order can be paid:
POST /orders/{id}/reserve runs before
POST /orders/{id}/pay. Payment must complete before
the order can ship: POST /orders/{id}/pay runs
before POST /orders/{id}/ship.
Accepted formats are PDF, DOCX, Markdown, YAML, JSON, and plain text, up to 2 MB per document. If you already keep an API handbook, an onboarding guide, or a wiki page, you can upload it as-is. The engine retrieves only the parts that matter for each endpoint and ignores the rest.
How our agent resolves a value from a knowledge document
Everything above is what you write. This is what turns it into a value the testing agent actually sends. The endpoint that needs sample_type cannot be talked into accepting anything by a cleverer prompt alone; it needs the literal string BLOOD, and that string has to travel from your document into that one request at the moment the agent is building it, unparaphrased, unsummarized, character-for-character correct.
A document is not pasted wholesale into a prompt: at 2 MB that would blow any context window, and most of it would be irrelevant to any single endpoint anyway. Instead, it runs through a pipeline that turns it into retrievable, embedded chunks, and at scan time the testing agent pulls only the chunks relevant to the endpoint it is currently trying to resolve. The pipeline has two phases:
- Ingestion phase that runs once per document
- Retrieval phase that runs during the scan
At a high level, with the exact stages and their parameters left out, it looks like this:

Ingestion and structure-aware chunking
Parsing and chunking are handled by a dedicated service. It first normalizes the uploaded document, then splits on structure before falling back to a size-based split. Each chunk comes back as two strings:
- A
textfield, which is what eventually gets shown to the agent - An
embed_textfield, which is what gets embedded for search
Embedding and storage
Each embed_text is turned into a fixed-dimension vector and stored in Postgres using the pgvector extension, alongside the chunk’s text, in a table grouped per document.
Retrieval during the scan
When the testing agent works an operation, it builds a query from that operation’s signature. That query is embedded and compared against the stored chunks with pgvector’s cosine-distance operator (<=>), returning the nearest chunks.
This is vector similarity, not keyword search. There is no separate exact-match index for codes. Exact codes still survive because of how the chunks are built: a chunk keeps its heading context, the literal codes live inside it, and the matched chunk is handed to the agent verbatim rather than summarized.
The embed_text/text split is what makes this work. You can tune what a chunk matches on without changing what the agent eventually reads.
Injection
The retrieved chunks are concatenated, capped at a character budget, and placed into the agent’s context as a labeled section that frames the content as authoritative: values and codes straight from the API owner’s own documentation, to be used exactly as written rather than second-guessed.
The chunks themselves go in unchanged. When the document says the value is NIST-SRM1921, that is the string the agent sees and uses.
In pseudocode, the per-operation loop is small:
for op in scan.operations:
query = signature(op) # verb + path + summary + params
qvec = embed(query)
chunks = sql("""
SELECT text
FROM document_chunks
ORDER BY embeddings <=> :qvec -- cosine distance, nearest first
LIMIT k
""", qvec=qvec)
knowledge = join(chunks)[:budget]
prompt = render(template, KnowledgeDocument=knowledge) # verbatim
This is the whole answer to the opening example: the retriever finds the chunk stating the specimen taxonomy, that chunk still contains the literal string BLOOD, and the injected system prompt hands the agent that string rather than a summary of it.
Nothing upstream of this pipeline had to change for the agent to stop guessing. The missing piece was the value, and this is the path that delivers it.
Before and after: Resolving a value that the specification never lists
A knowledge document can significantly expand the endpoint surface that can actually be tested.
Some fields cannot be inferred. No amount of reasoning derives a password from a schema, and the same is true of a specimen enum defined in an internal system, a site code assembled by local convention, or a calibration reference that must match a row in a lookup table. The specification describes the shape of the value, but does not provide the environment-specific value required to complete the request.
Deterministic substitution can construct the requests that require no additional domain knowledge. The remaining endpoints stay unresolved because the values they need are absent from the specification. The requests themselves can be well-formed and still be rejected.
A knowledge document supplies those missing values. It can also define a specific workflow to test: if the user provides the steps of a business process, the system can follow that sequence and test the corresponding operations in context. This makes previously unreachable endpoints accessible and allows operations that change state to be tested.
Sequence hints: Encoding call order
Not every gap between an agent and an API is about values. Sometimes, every value in a request is valid, and the call still fails because it ran before something else that had to happen first.
The pay-before-reserve failure from the opening has this shape. The two endpoints share no parameter that the testing system can use to infer a dependency between them. That ordering exists only in the API’s internal state machine, and the only place it is usually written down is documentation a human wrote for other humans.
Sequence hints let a knowledge document carry that ordering the same way it carries a hidden code or a lookup table, using the same kind of sentence shown above. Turning that sentence into something the testing agent obeys happens in three steps, run once per scan, right after the dependency graph is built and before the first sequence is generated. At a high level:
Extraction
For every enabled knowledge document, the engine locates every passage that mentions an operation by name, merges overlapping passages so shared context is not sent twice, and sends the result to an LLM with instructions to return every “this must run before that” relationship it can find.
Extraction only trusts what the text states or clearly implies. It never invents an operation the document does not mention, and a document that yields nothing usable is skipped.
Resolution
The extracted relationships are expressed in whatever names the document used, so they get matched against the scan’s real operations. Anything that cannot be matched is dropped.
For every operation, resolution then computes its full set of required predecessors, not just the one named directly, so a chain described piecemeal across a document (A before B, B before C) is understood as C requiring both A and B.
Bias
Sequence generation deprioritizes an operation in proportion to how many of its requirements have not succeeded yet. An operation that was merely attempted, including one that failed because it ran too early, does not count as satisfying a requirement. This only ever suppresses priority, never boosts it beyond what the rest of the engine’s logic already allows.
As with everything else in this article, this is additive: a scan with no knowledge documents, or none that describe any ordering, behaves exactly as it did before this feature existed.
An order the API never reveals
Some workflows cannot be reconstructed from the API specification alone. An operation may return an identifier that several later operations consume, but that shared identifier says nothing about which of those operations must happen first. The requests can all be individually valid while only certain sequences are accepted by the application.
Without additional context, the testing agent may have no reliable signal for recovering that order. It can attempt a downstream operation before its prerequisite has succeeded and receive a rejection that says little about what should have happened first. The dependency exists in the application’s internal state, not in the request or response shapes available to the agent.
A knowledge document makes that hidden order explicit. When it describes which operations must precede others, those relationships are resolved against the API’s actual operations and used to guide sequence generation. The agent can then exercise dependent operations in the context of the workflow they belong to, rather than relying only on dependencies that can be inferred from the API itself.
This matters most for stateful business processes, where reaching a later operation depends on completing the steps before it. The knowledge document does not change the API or make an invalid request valid. It supplies the missing business context needed to reach application states that the specification alone does not reveal.
Why this matters
The endpoints an agent cannot reach are the ones that have never been security-tested, and that is where vulnerabilities survive longest. The OWASP API Security Top 10 (2023) makes the dependency explicit: the most dangerous categories, broken object-level authorization and abuse of sensitive business flows, require a valid request before you can test anything meaningful.
You cannot check whether user A can read user B’s data if the tool cannot build a request that returns anyone’s data. Palo Alto Networks reported that 92% of organizations with API security products still experienced incidents: the tools are running, but they are not reaching the endpoints that matter.
The more business-critical an endpoint is, the stricter its input rules and call order tend to be, and the less likely a generic agent will ever test it. A knowledge document inverts that. The hardest endpoints to reach become testable, because the values they require are written down once and retrieved automatically wherever they are relevant.
The practical starting point is wherever your scans already fail: the endpoints with the most rejected requests are where a document has the most immediate effect, and often it is as little as listing the valid codes for a few fields or supplying one lookup table.
Your most critical endpoints are the ones your agent cannot reach. Book a demo to see how Equixly uses your domain knowledge to test what other tools skip.
FAQs
What format should knowledge documents be in?
PDF, DOCX, Markdown, YAML, JSON, or plain text, up to 2 MB per document. If you have API documentation in Confluence or a wiki, export it and upload it.
Do I need to write one for every endpoint?
No. Most endpoints work without extra context: the agent handles standard parameters and common formats on its own. Documents matter for endpoints with proprietary codes, custom formats, or field dependencies that cannot be inferred from the specification. Start with the ones your scans already struggle with.
How does the agent know which part of my document applies to which endpoint?
It does not rely on you tagging anything. For each operation, the agent builds a query from the operation’s verb, path, summary, and parameters, embeds it, and retrieves the few most similar chunks of your document by vector similarity. Sections about unrelated topics rank too low to be retrieved, so a large handbook costs nothing in noise. Extra content is stored but not pulled in unless it is relevant.
Can a knowledge document also tell the agents what order to call things in?
Yes. If your API enforces an order that is not visible from the request and response shapes, a resource must be reserved before it can be paid for, a background check must complete before an employee record can be promoted, and so on. Write it down as plainly as you would write down a valid code: X runs before Y. The engine extracts these relationships with an LLM, resolves them against the scan’s operations, and uses them to bias which operation the testing agent tries next. It never forces an operation the rest of the engine has already ruled out; it only makes the correct order more likely to be tried first.
Can I scope a knowledge document to specific parts of my API?
Yes. A document can be tied to a specific authentication setting or to specific services. This is useful when different parts of your API have different rules, or when a value such as a facility access code only applies when testing under particular credentials.
Alessio Dalla Piazza
CTO & FOUNDER
Former Founder & CTO of CYS4, he embarked on active digital surveillance work in 2014, collaborating with global and local law enforcement to combat terrorism and organized crime. He designed and utilized advanced eavesdropping technologies, identifying Zero-days in products like Skype, VMware, Safari, Docker, and IBM WebSphere. In June 2016, he transitioned to a research role at an international firm, where he crafted tools for automated offensive security and vulnerability detection. He discovered multiple vulnerabilities that, if exploited, would grant complete control. His expertise served the banking, insurance, and industrial sectors through Red Team operations, Incident Management, and Advanced Training, enhancing client security.
Ion Farima
Security Software Engineer
After graduating with a bachelor’s degree in computer science from the University of Florence in 2020, Ion distinguished himself with his thesis project, developing a framework for Symbolic Execution on PHP software for vulnerability scanning. His engagement with cybersecurity started in 2018 when he joined the CyberChallenge IT program. After his first year as a participant, he transitioned to a role as a tutor, guiding young cybersecurity enthusiasts. In 2021, he embarked on a professional path with CYS4 as a cybersecurity analyst, where he deepened his skills in penetration testing, red team operations, and various offensive cybersecurity techniques. By 2022, he played a significant role in a project focused on the automatic penetration testing of APIs, then incorporated as Equixly.
Nicolò Piccoli
Security Software Engineer
Nicolò studied Computer Engineering at the University of Verona, where he has honed significant expertise in software engineering and cybersecurity, with a particular emphasis on API security, which is the focus of his Master's thesis and ongoing research. His technical background also includes WiFi sensing research, high-performance computing, and malware analysis. Cisco CCNA-certified, Nicolò is committed to developing cutting-edge secure technologies and fostering innovation in the field.
