Traditional fuzzers generate random inputs or mutate existing ones. LLM Fuzz CI flips the model: an agent reads your code, understands your test assertions, and generates adversarial inputs designed to violate them. The agent writes inputs, your existing test suite replays them, and only real assertion failures get reported. No static analysis warnings, no CVE spam, no false positives.
The tool runs as a GitHub Action. You mark tests with a decorator, set a spend limit, and the agent tries to break your assertions. Every reported bug is a real failure of your own tests.
The agent receives three things: your source code, the test you marked, and the assertion logic inside that test. It does not generate new assertions. It generates inputs that it believes will cause your existing assertions to fail.
The workflow:
@pytest.mark.llm_fuzz(budget_usd=5, params=["user_input"]) or the vitest equivalent
The agent does not run the tests itself. It writes inputs to a file, and your standard test runner (pytest or vitest) replays them. This means the agent cannot hallucinate a vulnerability. If the test passes, nothing gets reported.
The tool does not persist state between runs by default. Each CI execution starts fresh. The agent does not remember previous findings or learn from past failures.
This design choice avoids complexity but means the agent might rediscover the same bug multiple times. The authors suggest running it on code changes or on a schedule, not on every commit.
For teams that want persistence, you could:
None of this is built in. The tool is stateless by design.
The feedback loop is simple: if the test fails, it's a real bug. If the test passes, the input was not adversarial enough.
The agent does not classify severity or decide what counts as a vulnerability. It generates inputs, the test runs, and the assertion either holds or breaks. The developer wrote the assertion, so the developer defined what "correct" means.
This eliminates the false positive problem that plagues static analysis tools. A CVE scanner might flag a dependency as vulnerable even if your code never calls the vulnerable function. LLM Fuzz CI only reports inputs that actually break your tests.
The downside: if your tests have weak assertions, the agent will not find bugs. If you test status_code == 200 but do not check the response body, the agent might generate inputs that corrupt data without failing the test.
The tool runs as a GitHub Action. The agent phase and the test execution phase are separate steps.
Agent phase:
Test phase:
The agent does not execute arbitrary code. It writes JSON or Python data structures. The test runner loads those inputs and passes them to your functions.
Security boundaries:
If your tests do not handle untrusted inputs safely, this tool will find that problem. That is the point.
The tool is a GitHub Action. You add it to your workflow YAML:
- name: LLM Fuzz CI
uses: tiime-software/llm-fuzz-ci@v1
with:
budget_usd: 10
llm_provider: openai
api_key: ${{ secrets.OPENAI_API_KEY }}
The budget_usd parameter caps spending. The agent stops generating inputs when it hits the limit. This prevents runaway costs if the agent gets stuck in a loop or generates thousands of inputs.
Cost scales with:
The authors report spending $5-10 per run on typical codebases. For a daily CI run, that is $150-300 per month. For a run on every PR, costs scale with PR volume.
Agent generates useless inputs:
If the agent does not understand the code, it might generate inputs that do not exercise interesting code paths. The test passes, nothing gets reported, and you wasted API credits.
Mitigation: start with simple tests and check that the agent generates plausible inputs. If it generates random strings for a function that expects JSON, the agent did not understand the schema.
Agent generates too many inputs:
If the agent finds a bug on the first input, great. If it generates 1,000 inputs before finding a bug, you pay for 1,000 LLM calls.
Mitigation: set a low budget for initial runs. Increase it if the agent is not finding bugs.
Test assertions are too weak:
If your test only checks response.status == 200, the agent might generate inputs that corrupt data, leak secrets, or trigger race conditions without failing the test.
Mitigation: write stronger assertions. Check response bodies, database state, and side effects.
Agent finds a bug but the report is unclear:
The tool reports which input caused the failure, but it does not explain why the input is adversarial or how to fix the bug.
Mitigation: read the generated input and the test failure. The input is usually small enough to understand manually.
| Dimension | Traditional Fuzzer | LLM Fuzz CI |
|---|---|---|
| Input generation | Random or mutation-based | LLM reads code and generates targeted inputs |
| Coverage | High code coverage, low semantic coverage | Low code coverage, high semantic coverage |
| False positives | Crashes and timeouts (may not be bugs) | Zero (only reports assertion failures) |
| Cost | CPU time (cheap) | LLM API calls (expensive) |
| Setup | Requires instrumentation and corpus | Requires marking tests |
| Best for | Finding memory corruption, crashes | Finding logic bugs, business rule violations |
Traditional fuzzers excel at finding crashes and memory corruption. LLM Fuzz CI excels at finding logic bugs that violate business rules but do not crash.
Example: a traditional fuzzer might find a buffer overflow in a parser. LLM Fuzz CI might find that a discount code can be applied twice, or that a user can access another user's data by manipulating a query parameter.
Good fit:
Bad fit:
The tool is not a replacement for static analysis, dependency scanning, or traditional fuzzing. It is a complement. Use it to find bugs that other tools miss because they do not understand your business logic.
LLM Fuzz CI solves the false positive problem by only reporting inputs that fail your own tests. This makes it useful for finding logic bugs that static analysis tools miss. The cost is manageable for daily or per-PR runs on small to medium codebases.
Use it if you have a strong test suite and want to find adversarial inputs that violate your assertions. Avoid it if your tests are weak or if you need exhaustive code coverage. The tool is only as good as your tests.
The stateless design is both a strength and a limitation. It keeps the tool simple but means you might rediscover the same bugs. For teams that want persistence, you will need to build your own state management layer.
The biggest risk is cost runaway if the agent generates thousands of inputs without finding bugs. Start with low budgets and monitor spending.