AI is making software testing cheap. Quality judgment is becoming more valuable Gurtam's QA team found that AI test generation is becoming easier much faster than evaluating whether generated tests are useful, after building an internal assistant that pulls documentation from a task number and produces test cases for review. The company says it has not gained enough confidence to remove the human validation step, because the assistant can still produce large numbers of unnecessary tests while missing risks that depend on undocumented product knowledge. For years, one of the practical constraints in software testing was time. Writing test cases takes time, preparing data takes time, regression takes time and analyzing what happened after an incident can take even more time. AI is changing this very quickly. Today, a QA engineer can generate test data, write scripts, analyze logs or create dozens of test cases in seconds. From the outside, this looks like an obvious productivity win: if we can create more tests faster, quality should improve as well. Our experience shows that it is not that simple. At Gurtam, we have been experimenting with AI in QA for some time. We tried different models, worked on prompts, used AI for generating test cases from documentation and eventually built our own internal assistant for this purpose. The technology is already genuinely useful, but the more we use it, the clearer one limitation becomes: generating tests is becoming easier much faster than evaluating whether those tests are actually useful. This distinction matters, because volume and quality are not the same thing. Test generation seems like one of the most natural applications of generative AI. A task already has requirements; there is documentation and the model can use that context to produce positive scenarios, negative scenarios and boundary cases almost instantly. Recent research https://link.springer.com/article/10.1007/s10586-026-06021-z into LLM-based test generation reflects the same opportunity, while also showing that the quality and usefulness of generated tests remain highly dependent on context, evaluation methods and human expertise. In practice, we often see two extremes. If there is a lot of documentation, AI can generate a large number of cases, many of which add little value. They may look technically reasonable, but they do not necessarily improve coverage in a meaningful way. The QA engineer still has to review them, understand which cases matter and remove the rest. The opposite happens when documentation is incomplete. In that case, the generated tests are often too superficial. They cover what is explicitly written, but they miss relationships, dependencies and risks that are obvious only if you understand the product and the surrounding system. The output can still look convincing, but that does not mean it gives enough confidence in the quality of the feature. This is where AI exposes a familiar problem rather than eliminating it. A model can work only with the context available to it. If important product knowledge lives in people’s heads, in previous incidents or in an understanding of architecture that was never documented, the generated test cases will not magically contain that knowledge. To make test generation more practical, we built an internal assistant. The workflow is simple: a QA engineer gives it the number of a task from our tracking system, the assistant finds the relevant documentation, combines that information with the task and generates test cases. We also created an interface where these cases can be reviewed, edited and removed if they are unnecessary. Once the QA engineer is satisfied with the result, the final set can be sent directly to the system where we store our test cases. This already removes a noticeable amount of repetitive work, which is useful on its own. But naturally, once you automate this much, the next question appears: can we remove the review step as well? In theory, it sounds reasonable. A product team creates and documents a task, a developer starts working on it and the AI assistant automatically generates the test cases. By the time development is finished, everything required for testing is already prepared. We considered such a flow. So far, though, our experience has not given us enough confidence to run it without human validation. The issue is not that AI cannot generate tests. It clearly can. The issue is that it can still produce a large number of unnecessary tests and at the same time miss something important. That means the review stage is not simply cosmetic. It is still part of the actual quality work. This does not mean I am skeptical about using AI in QA. There are many areas where I think it already works very well. Test-data generation is an obvious example. If we need boundary values, test JSON objects or large amounts of synthetic data with different parameters, AI can prepare this much faster than a person would manually. Log analysis is another strong case. Sometimes we deal with very large volumes of logs that no engineer would realistically read line by line. AI can help identify what happened during a particular incident or at a specific point in time and narrow down the area we need to investigate. It is also useful for generating small scripts or applications that remove repetitive work. This can include scripts for load testing, data generation, file analysis, one-off automated tests or simple interfaces for test environments. The common pattern is that the engineer already understands what needs to be done. AI helps execute the task faster. In those scenarios, it works as a very strong assistant. Where I would be much more careful is when we move from execution to decision-making. I would not currently trust AI with the complete process of writing and maintaining automated tests without supervision. In our own experiments, we have seen how easily it can generate unnecessary tests while still missing important scenarios. The same applies to complex, non-linear business logic. A model can process requirements, but understanding what is risky for the business or what behavior is acceptable in a particular context often requires more than what is written in documentation. I would also not rely on AI alone for deciding whether a release is ready. Release quality is not simply the sum of passed test cases. You need to understand what has changed, where the risks are, what dependencies exist, what was not covered and what the possible business impact is if something goes wrong. Data is another concern. Sensitive information disclosure is already recognized https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/? as a specific risk in LLM applications, particularly when systems process personal data, confidential business information or credentials. I would therefore be especially cautious about giving external models access to information that an organization cannot afford to expose. For now, the simplest principle still works well: trust, but verify. If AI takes over more repetitive regression work, basic test generation or routine execution, that does not mean QA becomes less relevant. It creates time for the things teams often postpone because there is never enough capacity. Testing architecture is one example. Risk modelling is another. There are also non-functional testing, test infrastructure, shift-left practices and involvement in business decisions before they become development problems. This is one of the reasons I see the profession moving towards Quality Engineering. Traditional QA is often associated with checking whether the finished product behaves according to requirements. Quality Engineering is broader. It means trying to reduce the possibility of defects appearing in the first place by influencing processes, architecture and automation. This also changes what we expect from experienced QA engineers. A technically strong tester should be able to test almost any task in the project. A strong Senior QA needs to see the project as a whole: its architecture, dependencies, risks and connection to business decisions. In a mature team, such a person becomes a link between business and development, not just the person who verifies a feature at the end. AI makes this transition more visible because the mechanical part of the work is becoming cheaper. The value moves towards understanding what actually needs attention. There is another side to this shift. AI is not only becoming a tool for QA; it is also becoming something QA engineers will increasingly have to test. We are already seeing more assistants, decision-making systems and AI-based product features. There are still no universally established practices for testing all of them, but I expect this to become a normal part of QA work. The direction is already visible in emerging AI assurance frameworks: NIST, for example, is developing https://www.nist.gov/artificial-intelligence/ai-research/tevv-athlon-framework-evaluating-ai-systems? structured approaches to test, evaluation, verification and validation of AI systems across LLMs, multimodal models and AI agents. Validating hallucinations, checking for toxicity, testing data leakage and resistance to prompt injection may soon become standard requirements in some QA roles. This makes the idea that AI will simply remove the need for QA difficult to take seriously. The more software relies on AI, the more new quality risks appear. AI models can make mistakes, hide errors, adjust tests to incorrect behavior or generate code with bugs. Another AI model may then fail to identify those same problems. Giving such a system full responsibility for release acceptance without human control is hard to imagine in any serious production environment today. It can. The more useful question is what happens after generation. Which cases are actually valuable? Which ones are redundant? What is missing? Which risks are not visible in the documentation? And does the final set of tests give us enough confidence in the product? AI is making the production of testing artefacts much faster. What it has not removed is the need for judgement. For QA engineers, that means the profession is not disappearing. It is becoming more analytical, more engineering-oriented and more connected to product and business decisions. And at least for now, the final responsibility for understanding whether the system is actually good enough still belongs to a person.