Dependencies & supply chain
Known-vulnerable packages, whether the affected code is actually reachable from your entry points, and a signed software bill of materials you can attach to a release.
Most scanners hand you a queue. Yours fills with rule matches nobody has confirmed, and the real issue sits on page four behind forty things that were never exploitable. Verdict Runtime is built the other way around: a finding has to earn its severity before it reaches you, it tells you exactly what evidence it has — and where we can fix it, we hand you the fix with proof that it worked.
' AND 1=1-- returned 24 rows, ' AND 1=2-- returned 0Illustrative example of the report format, not a real customer finding.
Nothing is silently promoted. Each finding carries the reason it holds the rank it does, so triage starts from evidence instead of a rule identifier.
Reported, ranked below corroborated findings, and marked as uncorroborated. Useful signal, presented as exactly that rather than as a confirmed vulnerability.
Dependencies, source code and infrastructure are each covered by two engines built by different teams, and findings are correlated across them. Agreement between two independent implementations is a materially stronger signal than either one alone, and it is the single most effective filter we have against false positives. Where only one engine exists for a class, its findings stay labelled as single-engine.
For the classes where a safe proof exists — injection, cross-site scripting, open redirect, XXE, command injection — a hand-written, reviewed proof runs against your deployed application inside a network-isolated sandbox whose egress is locked to your target alone. If it does not reproduce, it does not get called confirmed.
Static analysis, dependencies, infrastructure, the running application and your AI features all land in a single schema with consistent severity, location and confidence, then get deduplicated and correlated — so you read one report, not six.
Known-vulnerable packages, whether the affected code is actually reachable from your entry points, and a signed software bill of materials you can attach to a release.
Injection, authentication, access-control and cryptography defects, cross-correlated between two engines before anything is ranked high.
Committed credentials, cloud and Kubernetes misconfiguration, and build-pipeline weaknesses that let a workflow be hijacked.
Authenticated dynamic testing against a deployed environment, including broken object-level and function-level authorisation that only appears when two real accounts are compared.
Live adversarial testing of the model-backed parts of your product, covering the full OWASP Top 10 for LLM Applications plus emerging agentic threats.
Fixes prepared as reviewable patches, each re-scanned and run against your own test suite before it reaches you.
Most security tooling has no model of what an LLM feature can be talked into doing. These checks send real adversarial traffic at your live endpoint and judge the response against something concrete — a planted canary, a marker string, a tool name that should never have been reachable — rather than asking a model whether the answer looked unsafe.
| OWASP | Check | What a finding proves |
|---|---|---|
| LLM01 | Prompt injection | A crafted instruction overrode the system prompt and returned an exact attacker-chosen marker. |
| LLM01 | Indirect prompt injection | An instruction planted in fetched content — not in the user's message — was followed. |
| LLM01 | Multi-turn escalation | Defences that hold on a single message fail across a real multi-turn conversation. |
| LLM02 | Sensitive information disclosure | A canary secret planted in your own data came back in a model response. |
| LLM03 | Plugin and tool enumeration | The model disclosed the real integrations it can reach when asked to list them. |
| LLM04 | Persistent data poisoning | Content written by one account changed what a different account was told, on a fresh session. |
| LLM05 | Improper output handling | Model output reached the page or a downstream system without escaping. |
| LLM06 | Excessive agency | The model invoked a sensitive tool outside the user's stated intent. |
| LLM07 | System prompt leakage | A known verbatim fragment of your own system prompt was recovered. |
| LLM08 | Cross-tenant retrieval leakage | Retrieval crossed a tenant boundary and returned another tenant's canary. |
| LLM08 | Knowledge-base poisoning | A document submitted through your own product changed later answers after sync. |
| LLM09 | Hallucination under pressure | The model produced confident technical detail about an entity that does not exist. |
| LLM10 | Unbounded consumption | Oversized input and long conversations were accepted with no effective cap. |
| Agentic | Agent code execution | The agent ran an attacker-supplied command through its own execution tool. |
| Agentic | Tool parameter injection | An authorised tool was re-invoked with an out-of-scope parameter value. |
| Agentic | Agent-to-agent trust | A message claiming another agent's identity was accepted without verification. |
| Agentic | Identity and guardrail erosion | Sustained reframing across turns moved the model off its assigned identity and limits. |
Seventeen of twenty-one checks shown. Every one requires your written authorisation before it sends a single request.
Some flaws are invisible to any rule because nothing about them is syntactically wrong. A coupon that can be redeemed twice, a cart that accepts a negative quantity, a checkout that can be replayed at yesterday's price — the code is fine, the logic is not.
A real browser session walks your product the way a user does, capturing every API call the pages actually make — including the ones absent from your documentation. Captured calls are linked into a dependency graph, so a request needing a real order identifier gets one that genuinely exists.
A model is given the real captured traffic and asked what the business rules appear to be and how they might be broken. It sends real requests, reads the real responses, and adapts — a loop closer to a tester probing your app than to a scanner replaying a payload list.
You can also state invariants directly — a total must equal quantity times unit price, a balance must never go negative — and have them checked against live responses. A violation is a confirmed finding, not a judgement call.
Model-driven results are marked as needing review rather than promoted to confirmed. A machine that suggests a lead is useful; a machine that reports its own guesses as facts is the reason your queue is full.
Reporting a problem is the easy half. For dependency, code, infrastructure and pipeline findings, we prepare the change and then prove it — because an unverified fix is just another claim.
A dependency is moved to the lowest version that actually resolves the issue, with the upstream changelog summarised so you can see what else moves with it. Code, infrastructure and workflow fixes come as a diff against the real file, scoped to the finding.
The same analysis runs again against the patched tree. The finding has to be gone, and no new finding may appear in its place. This step is deterministic — no model is asked whether the fix looks correct.
Your own test suite runs against the patched code. If a fix turns a vulnerability into an outage, that is not a fix, and you find out from us rather than from production.
Nothing is committed, pushed or merged on your behalf. Every change arrives as a reviewable diff or a pull request you open yourself, with the verification results attached — and any fix we could not verify is handed over labelled as exactly that.
Every agent that can send a single packet at a target checks it against a scope file you write and sign off. The check is compiled into the tool, not a setting — no scope entry means the run exits before it touches a key or a report. There is no permissive mode.
The same standard applied to the product is applied to its own claims. We test this tool by fixing real security bugs in open-source projects that are not our customers, owe us nothing, and review the work in public. Four commitments we hold ourselves to.
Every application is different, so we quote rather than publish a rate card. You get the quote before any work starts, and the first assessment is free.
How many repositories and applications are in scope, whether live and logged-in testing is needed, and whether your product has AI features to test.
The report with evidence on every finding, the fixes we can prepare with their verification results, and a written note on anything we could not check.
One repository, one report, no obligation. You will get the findings, the evidence behind each one, the fixes we can prepare, and an honest note on what could not be checked.