ChatGPT’s New EU Status Exposes an AI Architecture Problem The European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act on Aug 31, 2026, citing its hybrid prompt-and-search nature and 159.1 million average monthly EU users, well above the 45 million threshold. Anton Fokin, CEO of Qtim, argues the classification exposes an architectural problem: once a system selects sources for users, the final answer is no longer sufficient evidence, so teams need request-level traceability joining model versions, tool calls, generated queries, source identifiers, policy decisions and rollbacks. He points to a 2026 Microsoft Research preprint analyzing 234,839 public ChatGPT conversations in which 79% of user inputs were classified as difficult to answer via conventional web search. A product team ships a “search the web” toggle behind a chatbot. The interface barely changes. The system now decides when to browse, rewrites the user’s question, selects sources, and compresses them into one answer. That small toggle creates a much harder audit problem. On Aug 31, 2026, the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act. The Commission called it a hybrid service because it answers prompts and can search the web. The classification followed what the product does beneath the interface. I’m Anton Fokin, CEO of Qtim. We build AI chatbots, RAG systems, and language-model integrations. The useful lesson for product teams is architectural: once a system selects information for users, the final answer stops being enough evidence. ChatGPT reported 159.1 million average monthly users in the EU in the official list, well above the 45 million threshold for a VLOSE. After notification, the service has four months to comply with the additional obligations for the largest search engines. Those include systemic-risk assessment and mitigation, annual independent audits, and data access under defined procedures. A 2026 Microsoft Research preprint gives the product question more weight. The authors examined 234,839 public ChatGPT conversations collected from 2023 through 2025. They classified 79% of user inputs as difficult to answer through conventional web search. For comparable searchable questions, ChatGPT responses covered less diverse information than Google results across most topics. The dataset is observational and does not represent every user. It still shows that a system can accept a broad question and return a narrower range of information. The designation imposes obligations on the named services. It does not automatically place every AI product under the DSA. The regulation covers intermediary services offered to recipients in the EU, regardless of the provider’s location. Outside the DSA, public AI products still need traceability when they choose sources for users. Before an AI feature ships, we ask: The answers define the product boundary more accurately than the word “chatbot.” A closed knowledge base may only need document identifiers and versioning. An open-web product needs a decision log, an evaluation pipeline, and a tested way to stop the feature. A generative-search answer is the end of at least four decisions: whether to search, which queries to send, which documents and passages to use, and which policies or filters to apply. Two identical answers can come from different sources. Two different answers can come from the same model version after web results or an index changes. We use a common request identifier to join the evidence chain. It should connect the model and instruction versions, tool calls, generated queries, source or chunk identifiers, policy decisions, the final answer, citations, retries, human intervention, and rollback events. The tempting shortcut is to keep every prompt, retrieved page, and response forever. That improves reproduction and creates a much larger privacy and security problem. Retention, masking, deletion, and access controls belong in the traceability design. A useful record reconstructs a decision without cloning the conversation database. Testing also moves one layer down. A fixed “golden answer” breaks as wording changes. More stable checks ask whether the right sources were eligible, the expected tool ran, a policy fired, or a high-risk condition stopped the workflow. Article 34 of the DSA requires systemic-risk assessments at least annually and before features likely to have a critical impact on identified risks. Product teams can translate that into a release gate. Four objects need to stay connected: the user scenario and the groups it affects; the risk and the observable signal that would reveal it; the mitigation, such as a limit, review step, interface change, or policy; the rollout plan, including monitoring, stop conditions, and rollback. The DSA does not prescribe a database schema or observability stack. This mapping is our engineering interpretation. Without it, the risk assessment sits in a document, the launch decision in a tracker, and runtime events across several logging systems. An audit becomes a reconstruction project. A risk record should point to a release version. A metric should point to its definition and dashboard. A launch decision needs an accountable person or role. A feature flag and a tested rollback path carry more weight than a promise to “switch it off quickly.” Very large platforms and search engines must provide data necessary for supervision to the Commission or the relevant Digital Services Coordinator after a reasoned request. Vetted researchers use a separate Article 40 process to obtain data for studying systemic risks. Direct access to production databases creates risks for privacy, trade secrets, and service stability. Ad hoc exports fail more quietly: fields change, metric definitions drift, and two datasets with the same label stop meaning the same thing. A usable access layer needs versioned schemas, a data dictionary, de-identification rules, access logs, and a reproducible sampling procedure. It also needs a boundary between evidence required for review and information whose disclosure would create another risk. Auditability costs time and infrastructure. Risk records, evaluation suites, instruction versioning, richer logs, and controlled exports add work before release. Building the same stack for a small internal assistant wastes time. We scale controls with the breadth of information the system can reach and the consequence of its answer. Our rule is simple: once AI selects sources and turns them into one answer, retrieval and generation should be designed as a single auditable decision chain. We build AI chatbots, RAG systems, and language-model integrations. In an AI product review https://qtim.pro/services/ai-chatbots?utm source=devto&utm medium=article&utm campaign=chatgpt-vlose-product-architecture&utm content=cta-bottom , we can identify the level of control a use case needs and design it into the architecture before the product scales.