# Search Journey Optimization with Gemini: From Query Fan-Out to Grounded Decisions

> Source: <https://dev.to/hjangid/search-journey-optimization-with-gemini-from-query-fan-out-to-grounded-decisions-3f99>
> Published: 2026-09-28 21:29:53+00:00

Building an explainable search-to-decision workflow with Gemini, Google Places, Google Search grounding, deterministic ranking, and Google Cloud.

**Open source**: [ai-search-journey-lab](https://github.com/hastimal/ai-search-journey-lab)

It started with one search query

I started `ai-search-journey-lab` with a question that looked simple enough:

Find a coffee shop near Geekdom in San Antonio for six people, quiet enough to work, open after 8 PM, and recommend the top three.

At first, I thought of it as a local-search problem.

Then I started implementing it.

Almost immediately, that single sentence turned into several separate engineering problems.

I needed to understand what the user actually meant by “**near Geekdom**.” I needed to capture the group size. I needed opening hours. I needed to find places that might actually work for six people sitting together.

Then there was the phrase:

quiet enough to work

That is where things became more interesting.

A structured Places response can give me things like an address, rating, hours, and a Place ID.

But “quiet enough to work” is not simply another field I can request.

So now I needed both structured place data and additional evidence.

Then, once I started generating more than one retrieval query, another problem appeared: ***the same place could be returned by multiple searches***.

That meant I also needed:

**Oh gosh..It's getting a lot!!**

At that point, I wasn't building a “***send a prompt to Gemini and print the answer***” application anymore.

I was building a **search journey**. (and later think about *search journey optimization!*)

And that became the idea behind **AI Search Journey Lab**.

This is the first article in a seven-part series built from the same open-source repository.

I wanted the series to follow the way the project itself evolved instead of creating seven unrelated AI demos.

The topics are also centered around the problems I am most interested in as a Google Cloud and AI developer: grounding, search behavior, data, agents, observability, and production deployment.

The progression looks roughly like this:

**search journey → grounded retrieval → visibility measurement → agentic analytics → tracing → AgentOps → healthcare application**

My goal with the series is not just to show what Gemini can generate.

I want to show how the pieces around the model matter just as much: APIs, deterministic logic, persistence, tools, traces, tests, and deployment.

Let's go back to the original request:

From an application perspective, that contains several different constraints:

And not all of those constraints should be handled the same way.

**For example:**

That distinction ended up shaping most of the architecture.

One decision influenced almost everything that came afterward:

I did not want Gemini to own the entire workflow.

It would have been very easy to send the whole prompt to a model and ask:

Find the best three places and explain why.

That would certainly produce an answer.

But it would make several things harder for me:

So I divided the workflow into three responsibilities.

**Gemini reasoning**

I use Gemini where natural-language interpretation is valuable:

**Retrieval**

I use external data sources for evidence:

**Deterministic application logic**

I keep things such as these in normal application code:

That gave me a much more inspectable pipeline.

I normally start the local environment with:

```
python scripts/run_app_locally.py
```

As the project evolved, I added observability components as well, so the local startup script eventually became responsible for more than just Streamlit.

For this article, though, I am concentrating on two parts:

The later articles will move into visibility, agents, and observability.

Before I search for anything, I need to understand what the request actually contains.

Conceptually, the Geekdom query becomes something like:

```
{
  "category": "coffee_shop",
  "location_reference": "Geekdom, San Antonio",
  "party_size": 6,
  "open_after": "20:00",
  "preferences": [
    "quiet",
    "work-friendly"
  ],
  "result_count": 3
}
```

The important part here is not the **JSON**.

It is the boundary between **language** and **application state**.

People do not normally talk in schemas.

A user might say:

Somewhere close to Geekdom. We have six people, don't want somewhere too loud, and we'll probably stay late.

The application still needs to understand:

```
{
  "party_size": 6,
  "noise_preference": "quiet",
  "minimum_close_time": "20:00"
}
```

That is a good job for **Gemini**.

The model interprets the language.

My application controls the structure.

Since I am more into development and technical architecture role, I prefer moving the result into a typed object rather than passing free-form model text deeper into the workflow.

A simplified version looks like this:

``` python
from pydantic import BaseModel, Field

class SearchIntent(BaseModel):
    category: str
    location_reference: str
    party_size: int | None = None
    open_after: str | None = None
    preferences: list[str] = Field(default_factory=list)
    result_count: int = 3
```

That gives me a cleaner boundary:

**Natural-language request → Gemini interpretation → Structured application state**

From this point forward, the rest of the pipeline has something predictable to work with.

I didn't want intent extraction buried somewhere inside one large function.

I gave the stage its own observable boundary.

Conceptually:

```
with trace_span(
    "gemini.extract_intent",
    attributes={"workflow.stage": "v1"},
):
    intent = extract_intent(user_query)
```

That may look like a small detail here.

It became much more useful later when I added OpenTelemetry.

A lesson I learned while building this project was:

Good architecture boundaries often become good observability boundaries later.

As you run locally, it should open in `http://localhost:8502/`

Let's see the journey now and wait to trace the results!

Once I knew what the user wanted, I ran into the next question:

What exactly should I search for?

Using the complete user sentence as one search query would throw many different requirements into the same retrieval call.

Instead, I started decomposing the request into a small retrieval plan.

A simplified fan-out might look like:

**Places task:**

coffee shops near Geekdom San Antonio

work-friendly coffee shops near downtown San Antonio

**Search evidence task:**

coffee near Geekdom quiet work open late

In the project, I treat that planning step separately:

```
with trace_span(
    "gemini.plan_fan_out",
    attributes={"workflow.stage": "v1"},
):
    fanout_plan = plan_fan_out(intent)
```

This is the part I refer to as query fan-out.

One user's decision question becomes multiple narrower retrieval tasks.

Initially, fan-out sounds like a simple idea:

If one query is useful, several queries should be better.

But that creates its own problems.

Imagine generating all of these:

```
coffee near Geekdom
quiet coffee near Geekdom
coffee for six near Geekdom
coffee with large tables near Geekdom
coffee open after 8 near Geekdom
coffee good for working near Geekdom
late-night coffee downtown San Antonio
```

Now I have:

The opposite extreme is also bad:

```
coffee near Geekdom
```

because most of the intent has disappeared.

So **query fan-out** became an optimization problem of its own:

Generate enough retrieval tasks to cover the user's important constraints without turning the workflow into uncontrolled query expansion.

That is a much more interesting problem than simply asking a model to generate ten related searches.

Once I had the fan-out plan, I needed actual entities.

For the local-search portion of the workflow, that meant Google Places API (New).

This gave me structured information such as:

A simplified retrieval call looks like this:

``` php
import requests

def search_places(query: str, api_key: str) -> list[dict]:
    endpoint = "https://places.googleapis.com/v1/places:searchText"

    payload = {
        "textQuery": query,
        "pageSize": 10,
    }

    headers = {
        "Content-Type": "application/json",
        "X-Goog-Api-Key": api_key,
        "X-Goog-FieldMask": ",".join(
            [
                "places.id",
                "places.displayName",
                "places.formattedAddress",
                "places.rating",
                "places.userRatingCount",
                "places.currentOpeningHours",
                "places.googleMapsUri",
            ]
        ),
    }

    response = requests.post(
        endpoint,
        json=payload,
        headers=headers,
        timeout=20,
    )

    response.raise_for_status()

    return response.json().get("places", [])
```

This is intentionally simplified for the article.

In the repository, there is additional application logic around the retrieval and candidate handling.

At first I was mostly interested in ratings, addresses, and opening hours.

Once fan-out was involved, identity became equally important.

Suppose two searches produce:

```
Merit Coffee
```

and

```
Merit Coffee — Southtown
```

String matching alone is not a great basis for deciding whether I have one business or two.

Using a canonical identifier such as `Place ID` gives me a stronger entity boundary.

That became important almost immediately when I started combining results from multiple retrieval tasks.

This is where the system became more than a Places demo.

The API can tell me a lot about a place.

But look again at part of the original requirement:

That is not the same kind of data as:

```
rating = 4.6
```

or:

```
open until 10 PM
```

The same problem appears in other searches:

good for a large student group

strong vegan options

suitable for someone with a specific preference

These are softer constraints.

I needed another evidence path.

Rather than asking Gemini:

Is this place good for working?

I wanted the system to investigate specific claims.

```
Candidate: Example Coffee

Verify:
- late-hour availability
- evidence relevant to working/studying
- evidence relevant to group seating

For each constraint:
- supported
- unsupported
- evidence
- citation
```

Conceptually, that stage looks like:

```
with trace_span(
    "search_grounding.verify_evidence",
    attributes={"workflow.stage": "v1"},
):
    evidence = verify_candidates(
        candidates=candidates,
        intent=intent,
    )
```

This gave me two different types of evidence:

**Google Places + Search-grounded evidence**

I wanted to preserve that distinction instead of blending everything together.

If operating hours come from Places, I should know that.

If “**work-friendly**” comes from **grounded Search evidence**, I should know that too.

And if I cannot verify something, I want the system to keep that state visible.

Once I started running multiple Places tasks, something predictable happened:

```
the same business started appearing more than once.
Places task A ───────┐
                     ├── Candidate X
Places task B ───────┘
```

If I ranked those raw results directly, Candidate X could appear stronger simply because more than one retrieval task discovered it.

That is **not** **relevance**.

That is **duplication**.

So I introduced normalization and deduplication before scoring.

A simplified form is:

``` php
def deduplicate_places(candidates: list[dict]) -> list[dict]:
    unique: dict[str, dict] = {}

    for candidate in candidates:
        place_id = candidate["place_id"]

        if place_id not in unique:
            unique[place_id] = candidate
            continue

        unique[place_id] = merge_candidate(
            unique[place_id],
            candidate,
        )

    return list(unique.values())
```

The sequence matters:

**retrieve → normalize → deduplicate → score**

I specifically did **NOT** want:

**retrieve → score duplicates → fix identity afterward**

For Journey Analysis, I used another demo query:

Find an Indian restaurant near Trinity University for eight students, open after 9 PM, with vegetarian options. Recommend the top three.

One of my representative executions produced:

```
Places tasks                     2
Raw Places results              20
Unique Places candidates        13

Search tasks                     2

Candidates with Search evidence  3
Places-only candidates          10
Unmatched evidence               8
```

That was a useful moment in the project.

**Instead of asking only:**

What are the top three restaurants?

**I could now ask:**

What actually happened during retrieval?

That became the purpose of **V2 — Journey Analysis**.

**Step 7: Retrieval and ranking are not the same thing**

This sounds obvious, but it became an important rule in my implementation.

**Retrieval asks:**

Which candidates might be relevant?

**Ranking asks:**

Which of those candidates best satisfies the original constraints?

A candidate should not rank higher simply because *two fan-out queries* happened to retrieve it.

And the **highest-rated business should not automatically win** if it fails an important constraint such as **opening hours**.

So after deduplication, I aggregate the evidence and score candidates explicitly.

This was another decision I made deliberately.

I did not want this architecture:

```
Retrieve 10 places
       ↓
Send all 10 to Gemini
       ↓
"Pick the best three"
```

That would give me an answer.

But it would make one of the most important decisions in the system opaque.

Instead, I keep the ranking logic inspectable in application code.

Depending on the workflow, signals can include things like:

```
with trace_span(
    "evidence.aggregate_and_score",
    attributes={"workflow.stage": "v1"},
):
    ranked_candidates = aggregate_and_score(
        candidates=candidates,
        evidence=evidence,
        intent=intent,
    )
```

The important property here is not one universal scoring formula.

It is that I can inspect *why a candidate received its score*.

``` python
def calculate_haversine_distance_miles(
    lat1: float,
    lon1: float,
    lat2: float,
    lon2: float,
) -> float:
    """Calculate the great circle distance in miles between two latitude/longitude points."""
    r_earth = 3958.8  # Earth radius in miles
    d_lat = math.radians(lat2 - lat1)
    d_lon = math.radians(lon2 - lon1)
    a = (
        math.sin(d_lat / 2.0) ** 2
        + math.cos(math.radians(lat1))
        * math.cos(math.radians(lat2))
        * math.sin(d_lon / 2.0) ** 2
    )
    c = 2.0 * math.atan2(math.sqrt(a), math.sqrt(1.0 - a))
    return round(r_earth * c, 2)

def calculate_proximity_score(distance_miles: float) -> float:
    """Calculate smooth bounded proximity bonus from distance in miles (max 12.0 pts).

    Formula: bonus = MAX_PROXIMITY_BONUS / (1.0 + distance_miles)
    - 0.0 mi  -> 12.00 pts
    - 0.5 mi  -> 8.00 pts
    - 1.0 mi  -> 6.00 pts
    - 2.0 mi  -> 4.00 pts
    - 5.0 mi  -> 2.00 pts
    - 190 mi  -> 0.06 pts (Houston vs San Antonio)
    """
    if distance_miles < 0.0:
        return 0.0
    bonus = MAX_PROXIMITY_BONUS / (1.0 + distance_miles)
    return round(bonus, 2)
```

[View the complete scoring implementation on GitHub](https://github.com/hastimal/ai-search-journey-lab/blob/main/src/ai_search_journey/ranking.py/)

Candidates are ranked after normalization and evidence aggregation rather than being silently reordered by the model.

Gemini still has an important job after the deterministic ranking stage.

I use it to turn the structured result into an answer that is useful to a person.

```
with trace_span(
    "gemini.synthesize_recommendations",
    attributes={"workflow.stage": "v1"},
):
    response = synthesize_recommendations(
        intent=intent,
        ranked_candidates=ranked_candidates,
    )
```

But notice what has changed.

At the start of the workflow, Gemini had:

```
one natural-language request
```

At the end, it can work with:

```
structured intent
       +
normalized candidates
       +
Places evidence
       +
Search evidence
       +
constraint states
       +
candidate scores
       +
final ranking
```

So instead of asking Gemini to discover, verify, rank, and explain everything at once, I give it a much narrower final job:

Explain a decision whose evidence has already been assembled.

That is the architecture I wanted.

At a high level, the workflow became:

And internally, I started giving the major stages explicit names such as:

```
gemini.extract_intent
gemini.plan_fan_out
places.text_search
search_grounding.verify_evidence
candidate.normalize_and_dedup
evidence.aggregate_and_score
gemini.synthesize_recommendations
```

Those names become important again later in this series when I start tracing the workflows.

Another goal for this project was to avoid treating AI code as somehow exempt from normal software-engineering practices.

After making changes, I run:

```
./.venv/bin/ruff check .
./.venv/bin/mypy src
./.venv/bin/pytest -q
```

I also wanted the project to go beyond my laptop.

The application runs on Google Cloud Run.

I can inspect the deployed service from the CLI:

```
gcloud run services describe ai-search-journey-lab \
  --region us-central1 \
  --format="yaml(
      metadata.name,
      status.url,
      status.latestReadyRevisionName,
      status.conditions
  )"
```

And verify the Streamlit health endpoint:

```
curl -fsS \
"https://ai-search-journey-lab-642110324230.us-central1.run.app/_stcore/health"
```

For me, this matters because it closes the development loop:

```
local code
   ↓
local execution
   ↓
tests
   ↓
container
   ↓
Cloud Run
   ↓
live application
```

The same Search-to-Decision workflow running from the Cloud Run deployment.

There is an important distinction I want to make.

This project is not:

I am using query fan-out as an application architecture pattern.

The project asks a developer-focused question:

If one natural-language request turns into multiple retrieval tasks, what engineering problems do I have to solve before I can return a grounded decision?

That is what I am trying to explore.

V1 and V2 focus on one user's decision.

But imagine running many related prompts.

Now I can start asking:

How often does a brand appear?

Which competitors are being mentioned?

Which sources are being cited?

Does the brand appear in the initial query, the fan-out queries, or only the final answer?

How does visibility change across different prompts?

That is where this project moved from Search Journey Optimization into AI Search Visibility.

And that is the next article.

Before I move fully into visibility analytics, Part 2 will go deeper into the actual local-search implementation:

Then **Part 3 will move into the AI Search Visibility & Brand Analyzer with Gemini and BigQuery**.

**Repository:** [GitHub Repository](https://github.com/hastimal/ai-search-journey-lab)

**Local startup:**

```
python scripts/run_app_locally.py
```


