Modernizing Unstructured Data Workflows: Alteryx Live Query meets Google Cloud BigQuery Alteryx One Live Query now integrates with Google Cloud BigQuery to let business users build unstructured-data pipelines in a browser without writing code, using secure SQL pushdown so transformation and reconciliation logic executes inside the warehouse. The workflow uses the Document Extract tool with user-selected models such as Gemini or Document AI to pull structured fields from invoice PDFs staged in Google Cloud Storage, then reconciles them against governed invoice history in BigQuery under fine-grained access policies. Alteryx says the pattern targets millions of unstructured documents, aiming for higher straight-through processing, faster reconciliation with fewer manual exceptions, stronger auditability, and reduced data movement. Alteryx One Live Query and Google Cloud BigQuery redefine how enterprises handle complex, unstructured data at scale by reducing reliance on disconnected tools, minimizing data movement, and enabling warehouse-native execution. Live Query provides an intuitive, browser-based environment that allows business users to build sophisticated data pipelines without writing a single line of code. When paired with the scale and performance of BigQuery, this duo enables secure SQL pushdown, meaning data transformation logic is executed directly within the warehouse rather than moving it across the network. This synergy is particularly potent for document-heavy workflows, where Live Query can orchestrate Google’s advanced AI and machine learning models — like Gemini — to extract intelligence from PDFs and images while maintaining the strict data governance and security of the BigQuery environment. Finance and operations teams often find themselves buried under a mountain of vendor invoices. These documents come as PDFs with inconsistent data, varying invoice number formats, and a high risk of manual error. Traditionally, data teams have had to resort to manual data extraction or use a fragmented series of tools to move data into a warehouse — a process that is both time-consuming and prone to duplicates or amount mismatches. With the integration of Live Query for BigQuery , you can now automate the entire lifecycle of an invoice. By processing PDFs, extracting key fields, and standardizing results directly within the BigQuery ecosystem, data teams can identify exceptions like anomalies or data issues within the documents faster while ensuring their governed enterprise data configured through the BigQuery Ffine grained access policies never has to leave the warehouse. Let’s dive deeper into how the Live Query solution works. Invoice PDFs arrive as unstructured documents in Google Cloud Storage. In the Live Query workflow, the Document Extract tool extracts structured fields using the model selected by the user, such as Gemini or Document AI. The results are then standardized for reliable matching and compared against historical invoice records already stored and governed in BigQuery. This automated workflow is designed to move from document to decision with minimal manual intervention. At enterprise scale, that pattern is not just about processing a handful of PDFs — it is about enabling analytics and AI over millions of unstructured documents while keeping reconciliation logic aligned with BigQuery. By combining AI-driven extraction with BigQuery-native execution, the solution provides: Higher straight-through processing : Reducing manual handling and allowing more invoices to move through the process automatically. Faster reconciliation and fewer manual exceptions: Comparing extracted invoices against BigQuery data earlier in the process so teams can identify duplicates and mismatches sooner, reducing the number of invoices that require manual review. Stronger auditability : Maintaining a clear trail of data from source to operational output. Reduced data movement : Keeping data resident in BigQuery and leverage its built-in capabilities for enhanced security, governance, and control. At a high level, invoice PDFs are staged in Cloud Storage, processed through a Live Query workflow in the browser, reconciled against governed invoice history already resident in BigQuery, and written back as operational outputs for finance teams. The key architectural advantage of this pattern is that workflow authoring happens in the browser, while transformation and reconciliation logic are executed through platform services against the governed data environment. The Live Query architecture is designed for warehouse-scale processing and governed execution. It uses secure SQL pushdown to execute transformation and reconciliation logic directly in BigQuery , and integrates with Gemini for advanced AI capabilities such as document extraction, classification, and summarization. Live Query Browser : Users author workflows visually in the browser, configuring document extraction, classification, cleansing, validation, and routing logic. Alteryx One Platform Services: The Transformation Service coordinates request handling, execution planning, and result retrieval between the browser workflow and the customer data environment. The services also add traceability and observability across workflow execution, strengthening both operational monitoring and auditability. Customer Google Cloud : Invoice PDFs are staged in Google Cloud Storage GCS , historical invoice data remains governed in BigQuery , AI-backed extraction runs against that environment, and the workflow outputs are written back into BigQuery. In this pattern, the browser is the authoring layer, Alteryx One Platform Services are the coordination layer, and Google Cloud is the governed data and execution layer. It allows workflow logic to be authored visually in Live Query while execution remains aligned with governed data and warehouse-scale processing in BigQuery . Let’s take a closer look at how the Live Query workflow handles unstructured PDFs, using an invoice as the example. Ingest invoice PDFs The workflow begins by pointing to a directory input in Google Cloud Storage, where the raw PDF files are stored. In practice, this step does more than just identify a folder: it enumerates the PDF locations and returns file-level metadata — such as absolute file paths, size, and created / updated timestamps — so the workflow can reliably pass the right document set into downstream extraction steps. Extract key fields Using Alteryx’s Document Extract tool, the workflow processes PDF invoices stored in Google Cloud Storage and extracts structured or semi-structured results directly in BigQuery. The extraction step uses the user-selected model, either Document AI or Gemini AI, depending on the workflow configuration. Under the hood, the tool builds a temporary external object table from the input document URIs and executes the appropriate function — ML.PROCESS DOCUMENT for Document AI model or AI.GENERATE for Gemini AI model. For this workflow, the Gemini path is used, with AI.GENERATE returning the requested invoice fields into the downstream Live Query workflow. During interactive design, Live Query limits processing to a small subset of document URIs for preview, while runtime execution processes the full set of invoice PDFs. Classify extracted content After extraction, the workflow uses Alteryx’s Classify tool to apply zero-shot classification to the extracted invoice text. In this use case, line-item content is labeled into predefined business categories — such as industrial or office supplies — using AI.CLASSIFY in BigQuery, which returns a new predicted-label column into the downstream workflow. This business context makes the extracted data more useful for downstream reconciliation, exception handling, and reporting. Normalize data The extracted invoice numbers are cleansed and standardized so that minor formatting differences — such as whitespace or inconsistent formatting — do not hide potential duplicates during reconciliation. Compare against history The workflow joins the normalized current invoice records against historical invoice data already resident in BigQuery. This reconciliation step identifies records with matching invoice numbers already present in history and flags them as potential duplicates. Flag exceptions A formula step then applies business-rule logic to the remaining records. Those rules can vary by workflow and may include checks for amount mismatches, missing required fields, invalid totals, or other finance-specific exceptions. In this example, one rule checks whether invoice amount matches net amount + tax amount and flags mismatches for review. Produce operational outputs The workflow produces two operational outputs: one for invoices matched to historical records and flagged as potential duplicates, and another for unmatched current invoices that have been validated, labeled for exceptions where needed, and prepared for unpaid-invoice processing. The result is not just extracted data, but governed operational outputs that finance teams can review and act on immediately. Our joint customers are excited for these new capabilities that will help improve business processes. “What stood out to me about Alteryx One: Google Edition was how naturally it fits into the Google Cloud experience. I'm excited about the opportunity to make analytics more accessible by enabling business users to work with BigQuery data through an intuitive, governed experience.” - Michael Wyant, Vice President, Enterprise Data & Corporate Solutions, Papa Johns This use case is more than just a simple extraction tool; it represents a strong Google Cloud-native pattern for transforming unstructured documents into actionable operational outputs. Instead of moving warehouse data into disconnected preparation tools, the workflow keeps storage, extraction, reconciliation, and output generation aligned with governed enterprise data already resident in BigQuery. By combining the intuitive design of Alteryx One Live Query with the scale and AI capabilities of BigQuery, finance teams can reduce manual effort, accelerate exception handling, and focus on higher-value strategic initiatives.