{"slug": "search-journey-optimization-with-gemini-from-query-fan-out-to-grounded-decisions", "title": "Search Journey Optimization with Gemini: From Query Fan-Out to Grounded Decisions", "summary": "A developer built ai-search-journey-lab, an open-source workflow that combines Gemini reasoning, Google Places and Google Search grounding, and deterministic application logic to turn a single natural-language request into an explainable search-to-decision pipeline. The project splits responsibilities so Gemini handles natural-language interpretation while retrieval and ranking stay in inspectable application code, and it is documented in a seven-part series covering grounding, visibility measurement, agents, tracing and deployment.", "body_md": "Building an explainable search-to-decision workflow with Gemini, Google Places, Google Search grounding, deterministic ranking, and Google Cloud.\n\n**Open source**: [ai-search-journey-lab](https://github.com/hastimal/ai-search-journey-lab)\n\nIt started with one search query\n\nI started `ai-search-journey-lab` with a question that looked simple enough:\n\nFind a coffee shop near Geekdom in San Antonio for six people, quiet enough to work, open after 8 PM, and recommend the top three.\n\nAt first, I thought of it as a local-search problem.\n\nThen I started implementing it.\n\nAlmost immediately, that single sentence turned into several separate engineering problems.\n\nI needed to understand what the user actually meant by “**near Geekdom**.” I needed to capture the group size. I needed opening hours. I needed to find places that might actually work for six people sitting together.\n\nThen there was the phrase:\n\nquiet enough to work\n\nThat is where things became more interesting.\n\nA structured Places response can give me things like an address, rating, hours, and a Place ID.\n\nBut “quiet enough to work” is not simply another field I can request.\n\nSo now I needed both structured place data and additional evidence.\n\nThen, once I started generating more than one retrieval query, another problem appeared: ***the same place could be returned by multiple searches***.\n\nThat meant I also needed:\n\n**Oh gosh..It's getting a lot!!**\n\nAt that point, I wasn't building a “***send a prompt to Gemini and print the answer***” application anymore.\n\nI was building a **search journey**. (and later think about *search journey optimization!*)\n\nAnd that became the idea behind **AI Search Journey Lab**.\n\nThis is the first article in a seven-part series built from the same open-source repository.\n\nI wanted the series to follow the way the project itself evolved instead of creating seven unrelated AI demos.\n\nThe topics are also centered around the problems I am most interested in as a Google Cloud and AI developer: grounding, search behavior, data, agents, observability, and production deployment.\n\nThe progression looks roughly like this:\n\n**search journey → grounded retrieval → visibility measurement → agentic analytics → tracing → AgentOps → healthcare application**\n\nMy goal with the series is not just to show what Gemini can generate.\n\nI want to show how the pieces around the model matter just as much: APIs, deterministic logic, persistence, tools, traces, tests, and deployment.\n\nLet's go back to the original request:\n\nFrom an application perspective, that contains several different constraints:\n\nAnd not all of those constraints should be handled the same way.\n\n**For example:**\n\nThat distinction ended up shaping most of the architecture.\n\nOne decision influenced almost everything that came afterward:\n\nI did not want Gemini to own the entire workflow.\n\nIt would have been very easy to send the whole prompt to a model and ask:\n\nFind the best three places and explain why.\n\nThat would certainly produce an answer.\n\nBut it would make several things harder for me:\n\nSo I divided the workflow into three responsibilities.\n\n**Gemini reasoning**\n\nI use Gemini where natural-language interpretation is valuable:\n\n**Retrieval**\n\nI use external data sources for evidence:\n\n**Deterministic application logic**\n\nI keep things such as these in normal application code:\n\nThat gave me a much more inspectable pipeline.\n\nI normally start the local environment with:\n\n```\npython scripts/run_app_locally.py\n```\n\nAs the project evolved, I added observability components as well, so the local startup script eventually became responsible for more than just Streamlit.\n\nFor this article, though, I am concentrating on two parts:\n\nThe later articles will move into visibility, agents, and observability.\n\nBefore I search for anything, I need to understand what the request actually contains.\n\nConceptually, the Geekdom query becomes something like:\n\n```\n{\n  \"category\": \"coffee_shop\",\n  \"location_reference\": \"Geekdom, San Antonio\",\n  \"party_size\": 6,\n  \"open_after\": \"20:00\",\n  \"preferences\": [\n    \"quiet\",\n    \"work-friendly\"\n  ],\n  \"result_count\": 3\n}\n```\n\nThe important part here is not the **JSON**.\n\nIt is the boundary between **language** and **application state**.\n\nPeople do not normally talk in schemas.\n\nA user might say:\n\nSomewhere close to Geekdom. We have six people, don't want somewhere too loud, and we'll probably stay late.\n\nThe application still needs to understand:\n\n```\n{\n  \"party_size\": 6,\n  \"noise_preference\": \"quiet\",\n  \"minimum_close_time\": \"20:00\"\n}\n```\n\nThat is a good job for **Gemini**.\n\nThe model interprets the language.\n\nMy application controls the structure.\n\nSince I am more into development and technical architecture role, I prefer moving the result into a typed object rather than passing free-form model text deeper into the workflow.\n\nA simplified version looks like this:\n\n``` python\nfrom pydantic import BaseModel, Field\n\nclass SearchIntent(BaseModel):\n    category: str\n    location_reference: str\n    party_size: int | None = None\n    open_after: str | None = None\n    preferences: list[str] = Field(default_factory=list)\n    result_count: int = 3\n```\n\nThat gives me a cleaner boundary:\n\n**Natural-language request → Gemini interpretation → Structured application state**\n\nFrom this point forward, the rest of the pipeline has something predictable to work with.\n\nI didn't want intent extraction buried somewhere inside one large function.\n\nI gave the stage its own observable boundary.\n\nConceptually:\n\n```\nwith trace_span(\n    \"gemini.extract_intent\",\n    attributes={\"workflow.stage\": \"v1\"},\n):\n    intent = extract_intent(user_query)\n```\n\nThat may look like a small detail here.\n\nIt became much more useful later when I added OpenTelemetry.\n\nA lesson I learned while building this project was:\n\nGood architecture boundaries often become good observability boundaries later.\n\nAs you run locally, it should open in `http://localhost:8502/`\n\nLet's see the journey now and wait to trace the results!\n\nOnce I knew what the user wanted, I ran into the next question:\n\nWhat exactly should I search for?\n\nUsing the complete user sentence as one search query would throw many different requirements into the same retrieval call.\n\nInstead, I started decomposing the request into a small retrieval plan.\n\nA simplified fan-out might look like:\n\n**Places task:**\n\ncoffee shops near Geekdom San Antonio\n\nwork-friendly coffee shops near downtown San Antonio\n\n**Search evidence task:**\n\ncoffee near Geekdom quiet work open late\n\nIn the project, I treat that planning step separately:\n\n```\nwith trace_span(\n    \"gemini.plan_fan_out\",\n    attributes={\"workflow.stage\": \"v1\"},\n):\n    fanout_plan = plan_fan_out(intent)\n```\n\nThis is the part I refer to as query fan-out.\n\nOne user's decision question becomes multiple narrower retrieval tasks.\n\nInitially, fan-out sounds like a simple idea:\n\nIf one query is useful, several queries should be better.\n\nBut that creates its own problems.\n\nImagine generating all of these:\n\n```\ncoffee near Geekdom\nquiet coffee near Geekdom\ncoffee for six near Geekdom\ncoffee with large tables near Geekdom\ncoffee open after 8 near Geekdom\ncoffee good for working near Geekdom\nlate-night coffee downtown San Antonio\n```\n\nNow I have:\n\nThe opposite extreme is also bad:\n\n```\ncoffee near Geekdom\n```\n\nbecause most of the intent has disappeared.\n\nSo **query fan-out** became an optimization problem of its own:\n\nGenerate enough retrieval tasks to cover the user's important constraints without turning the workflow into uncontrolled query expansion.\n\nThat is a much more interesting problem than simply asking a model to generate ten related searches.\n\nOnce I had the fan-out plan, I needed actual entities.\n\nFor the local-search portion of the workflow, that meant Google Places API (New).\n\nThis gave me structured information such as:\n\nA simplified retrieval call looks like this:\n\n``` php\nimport requests\n\ndef search_places(query: str, api_key: str) -> list[dict]:\n    endpoint = \"https://places.googleapis.com/v1/places:searchText\"\n\n    payload = {\n        \"textQuery\": query,\n        \"pageSize\": 10,\n    }\n\n    headers = {\n        \"Content-Type\": \"application/json\",\n        \"X-Goog-Api-Key\": api_key,\n        \"X-Goog-FieldMask\": \",\".join(\n            [\n                \"places.id\",\n                \"places.displayName\",\n                \"places.formattedAddress\",\n                \"places.rating\",\n                \"places.userRatingCount\",\n                \"places.currentOpeningHours\",\n                \"places.googleMapsUri\",\n            ]\n        ),\n    }\n\n    response = requests.post(\n        endpoint,\n        json=payload,\n        headers=headers,\n        timeout=20,\n    )\n\n    response.raise_for_status()\n\n    return response.json().get(\"places\", [])\n```\n\nThis is intentionally simplified for the article.\n\nIn the repository, there is additional application logic around the retrieval and candidate handling.\n\nAt first I was mostly interested in ratings, addresses, and opening hours.\n\nOnce fan-out was involved, identity became equally important.\n\nSuppose two searches produce:\n\n```\nMerit Coffee\n```\n\nand\n\n```\nMerit Coffee — Southtown\n```\n\nString matching alone is not a great basis for deciding whether I have one business or two.\n\nUsing a canonical identifier such as `Place ID` gives me a stronger entity boundary.\n\nThat became important almost immediately when I started combining results from multiple retrieval tasks.\n\nThis is where the system became more than a Places demo.\n\nThe API can tell me a lot about a place.\n\nBut look again at part of the original requirement:\n\nThat is not the same kind of data as:\n\n```\nrating = 4.6\n```\n\nor:\n\n```\nopen until 10 PM\n```\n\nThe same problem appears in other searches:\n\ngood for a large student group\n\nstrong vegan options\n\nsuitable for someone with a specific preference\n\nThese are softer constraints.\n\nI needed another evidence path.\n\nRather than asking Gemini:\n\nIs this place good for working?\n\nI wanted the system to investigate specific claims.\n\n```\nCandidate: Example Coffee\n\nVerify:\n- late-hour availability\n- evidence relevant to working/studying\n- evidence relevant to group seating\n\nFor each constraint:\n- supported\n- unsupported\n- evidence\n- citation\n```\n\nConceptually, that stage looks like:\n\n```\nwith trace_span(\n    \"search_grounding.verify_evidence\",\n    attributes={\"workflow.stage\": \"v1\"},\n):\n    evidence = verify_candidates(\n        candidates=candidates,\n        intent=intent,\n    )\n```\n\nThis gave me two different types of evidence:\n\n**Google Places + Search-grounded evidence**\n\nI wanted to preserve that distinction instead of blending everything together.\n\nIf operating hours come from Places, I should know that.\n\nIf “**work-friendly**” comes from **grounded Search evidence**, I should know that too.\n\nAnd if I cannot verify something, I want the system to keep that state visible.\n\nOnce I started running multiple Places tasks, something predictable happened:\n\n```\nthe same business started appearing more than once.\nPlaces task A ───────┐\n                     ├── Candidate X\nPlaces task B ───────┘\n```\n\nIf I ranked those raw results directly, Candidate X could appear stronger simply because more than one retrieval task discovered it.\n\nThat is **not** **relevance**.\n\nThat is **duplication**.\n\nSo I introduced normalization and deduplication before scoring.\n\nA simplified form is:\n\n``` php\ndef deduplicate_places(candidates: list[dict]) -> list[dict]:\n    unique: dict[str, dict] = {}\n\n    for candidate in candidates:\n        place_id = candidate[\"place_id\"]\n\n        if place_id not in unique:\n            unique[place_id] = candidate\n            continue\n\n        unique[place_id] = merge_candidate(\n            unique[place_id],\n            candidate,\n        )\n\n    return list(unique.values())\n```\n\nThe sequence matters:\n\n**retrieve → normalize → deduplicate → score**\n\nI specifically did **NOT** want:\n\n**retrieve → score duplicates → fix identity afterward**\n\nFor Journey Analysis, I used another demo query:\n\nFind an Indian restaurant near Trinity University for eight students, open after 9 PM, with vegetarian options. Recommend the top three.\n\nOne of my representative executions produced:\n\n```\nPlaces tasks                     2\nRaw Places results              20\nUnique Places candidates        13\n\nSearch tasks                     2\n\nCandidates with Search evidence  3\nPlaces-only candidates          10\nUnmatched evidence               8\n```\n\nThat was a useful moment in the project.\n\n**Instead of asking only:**\n\nWhat are the top three restaurants?\n\n**I could now ask:**\n\nWhat actually happened during retrieval?\n\nThat became the purpose of **V2 — Journey Analysis**.\n\n**Step 7: Retrieval and ranking are not the same thing**\n\nThis sounds obvious, but it became an important rule in my implementation.\n\n**Retrieval asks:**\n\nWhich candidates might be relevant?\n\n**Ranking asks:**\n\nWhich of those candidates best satisfies the original constraints?\n\nA candidate should not rank higher simply because *two fan-out queries* happened to retrieve it.\n\nAnd the **highest-rated business should not automatically win** if it fails an important constraint such as **opening hours**.\n\nSo after deduplication, I aggregate the evidence and score candidates explicitly.\n\nThis was another decision I made deliberately.\n\nI did not want this architecture:\n\n```\nRetrieve 10 places\n       ↓\nSend all 10 to Gemini\n       ↓\n\"Pick the best three\"\n```\n\nThat would give me an answer.\n\nBut it would make one of the most important decisions in the system opaque.\n\nInstead, I keep the ranking logic inspectable in application code.\n\nDepending on the workflow, signals can include things like:\n\n```\nwith trace_span(\n    \"evidence.aggregate_and_score\",\n    attributes={\"workflow.stage\": \"v1\"},\n):\n    ranked_candidates = aggregate_and_score(\n        candidates=candidates,\n        evidence=evidence,\n        intent=intent,\n    )\n```\n\nThe important property here is not one universal scoring formula.\n\nIt is that I can inspect *why a candidate received its score*.\n\n``` python\ndef calculate_haversine_distance_miles(\n    lat1: float,\n    lon1: float,\n    lat2: float,\n    lon2: float,\n) -> float:\n    \"\"\"Calculate the great circle distance in miles between two latitude/longitude points.\"\"\"\n    r_earth = 3958.8  # Earth radius in miles\n    d_lat = math.radians(lat2 - lat1)\n    d_lon = math.radians(lon2 - lon1)\n    a = (\n        math.sin(d_lat / 2.0) ** 2\n        + math.cos(math.radians(lat1))\n        * math.cos(math.radians(lat2))\n        * math.sin(d_lon / 2.0) ** 2\n    )\n    c = 2.0 * math.atan2(math.sqrt(a), math.sqrt(1.0 - a))\n    return round(r_earth * c, 2)\n\ndef calculate_proximity_score(distance_miles: float) -> float:\n    \"\"\"Calculate smooth bounded proximity bonus from distance in miles (max 12.0 pts).\n\n    Formula: bonus = MAX_PROXIMITY_BONUS / (1.0 + distance_miles)\n    - 0.0 mi  -> 12.00 pts\n    - 0.5 mi  -> 8.00 pts\n    - 1.0 mi  -> 6.00 pts\n    - 2.0 mi  -> 4.00 pts\n    - 5.0 mi  -> 2.00 pts\n    - 190 mi  -> 0.06 pts (Houston vs San Antonio)\n    \"\"\"\n    if distance_miles < 0.0:\n        return 0.0\n    bonus = MAX_PROXIMITY_BONUS / (1.0 + distance_miles)\n    return round(bonus, 2)\n```\n\n[View the complete scoring implementation on GitHub](https://github.com/hastimal/ai-search-journey-lab/blob/main/src/ai_search_journey/ranking.py/)\n\nCandidates are ranked after normalization and evidence aggregation rather than being silently reordered by the model.\n\nGemini still has an important job after the deterministic ranking stage.\n\nI use it to turn the structured result into an answer that is useful to a person.\n\n```\nwith trace_span(\n    \"gemini.synthesize_recommendations\",\n    attributes={\"workflow.stage\": \"v1\"},\n):\n    response = synthesize_recommendations(\n        intent=intent,\n        ranked_candidates=ranked_candidates,\n    )\n```\n\nBut notice what has changed.\n\nAt the start of the workflow, Gemini had:\n\n```\none natural-language request\n```\n\nAt the end, it can work with:\n\n```\nstructured intent\n       +\nnormalized candidates\n       +\nPlaces evidence\n       +\nSearch evidence\n       +\nconstraint states\n       +\ncandidate scores\n       +\nfinal ranking\n```\n\nSo instead of asking Gemini to discover, verify, rank, and explain everything at once, I give it a much narrower final job:\n\nExplain a decision whose evidence has already been assembled.\n\nThat is the architecture I wanted.\n\nAt a high level, the workflow became:\n\nAnd internally, I started giving the major stages explicit names such as:\n\n```\ngemini.extract_intent\ngemini.plan_fan_out\nplaces.text_search\nsearch_grounding.verify_evidence\ncandidate.normalize_and_dedup\nevidence.aggregate_and_score\ngemini.synthesize_recommendations\n```\n\nThose names become important again later in this series when I start tracing the workflows.\n\nAnother goal for this project was to avoid treating AI code as somehow exempt from normal software-engineering practices.\n\nAfter making changes, I run:\n\n```\n./.venv/bin/ruff check .\n./.venv/bin/mypy src\n./.venv/bin/pytest -q\n```\n\nI also wanted the project to go beyond my laptop.\n\nThe application runs on Google Cloud Run.\n\nI can inspect the deployed service from the CLI:\n\n```\ngcloud run services describe ai-search-journey-lab \\\n  --region us-central1 \\\n  --format=\"yaml(\n      metadata.name,\n      status.url,\n      status.latestReadyRevisionName,\n      status.conditions\n  )\"\n```\n\nAnd verify the Streamlit health endpoint:\n\n```\ncurl -fsS \\\n\"https://ai-search-journey-lab-642110324230.us-central1.run.app/_stcore/health\"\n```\n\nFor me, this matters because it closes the development loop:\n\n```\nlocal code\n   ↓\nlocal execution\n   ↓\ntests\n   ↓\ncontainer\n   ↓\nCloud Run\n   ↓\nlive application\n```\n\nThe same Search-to-Decision workflow running from the Cloud Run deployment.\n\nThere is an important distinction I want to make.\n\nThis project is not:\n\nI am using query fan-out as an application architecture pattern.\n\nThe project asks a developer-focused question:\n\nIf one natural-language request turns into multiple retrieval tasks, what engineering problems do I have to solve before I can return a grounded decision?\n\nThat is what I am trying to explore.\n\nV1 and V2 focus on one user's decision.\n\nBut imagine running many related prompts.\n\nNow I can start asking:\n\nHow often does a brand appear?\n\nWhich competitors are being mentioned?\n\nWhich sources are being cited?\n\nDoes the brand appear in the initial query, the fan-out queries, or only the final answer?\n\nHow does visibility change across different prompts?\n\nThat is where this project moved from Search Journey Optimization into AI Search Visibility.\n\nAnd that is the next article.\n\nBefore I move fully into visibility analytics, Part 2 will go deeper into the actual local-search implementation:\n\nThen **Part 3 will move into the AI Search Visibility & Brand Analyzer with Gemini and BigQuery**.\n\n**Repository:** [GitHub Repository](https://github.com/hastimal/ai-search-journey-lab)\n\n**Local startup:**\n\n```\npython scripts/run_app_locally.py\n```\n\n", "url": "https://wpnews.pro/news/search-journey-optimization-with-gemini-from-query-fan-out-to-grounded-decisions", "canonical_source": "https://dev.to/hjangid/search-journey-optimization-with-gemini-from-query-fan-out-to-grounded-decisions-3f99", "published_at": "2026-09-28 21:29:53+00:00", "updated_at": "2026-09-28 21:49:05.215409+00:00", "lang": "en", "topics": ["ai-search", "generative-engine-optimization", "ai-agents", "ai-tools", "developer-tools"], "entities": ["Gemini", "Google Places", "Google Search", "Google Cloud", "Geekdom", "ai-search-journey-lab"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/search-journey-optimization-with-gemini-from-query-fan-out-to-grounded-decisions", "markdown": "https://wpnews.pro/news/search-journey-optimization-with-gemini-from-query-fan-out-to-grounded-decisions.md", "text": "https://wpnews.pro/news/search-journey-optimization-with-gemini-from-query-fan-out-to-grounded-decisions.txt", "jsonld": "https://wpnews.pro/news/search-journey-optimization-with-gemini-from-query-fan-out-to-grounded-decisions.jsonld"}}