SatQuery AI: Building a Conversational System for Satellite-Image Analysis
I built SatQuery AI around a simple idea: I wanted users to interact with satellite and Earth-observation imagery through natural language rather than having to translate every question into a sequence of specialized image-processing and geospatial operations.
A user should be able to ask:
“Where has vegetation decreased?”
or:
“What changed between these two satellite images?”
or:
“Detect buildings in this region.”
But building a system that answers these questions reliably is fundamentally different from building a chatbot.
The distinction I kept coming back to was:
A language model can explain an answer, but the satellite-analysis pipeline has to provide the evidence.
That distinction shaped the architecture of SatQuery AI.
From Questions to Evidence A conventional conversational AI system can take a question, generate an answer, and return it directly. That approach is not sufficient when the answer depends on information contained in an image.
Suppose I ask:
“Show me the areas where vegetation decreased between these images.”
A language model can generate a plausible explanation of vegetation change. But generating a sentence is not the same as detecting vegetation change.
For SatQuery AI, the workflow is closer to: Ask → Understand → Analyze → Verify → Visualize → Explain
The natural-language layer first needs to understand what the user is asking. The system then needs to determine what type of Earth-observation analysis is appropriate and execute that analysis against the available imagery.
Depending on the request, that could involve object detection, segmentation, change detection, image comparison, vegetation analysis, land-use and land-cover analysis, object counting, or other geospatial operations.
The important part is that the analytical pipeline produces an actual result.
For the vegetation example, that result might contain regions identified as changed, measurements associated with those regions, percentages, confidence information, or other relevant geospatial information. Only after that evidence exists does the conversational layer have something meaningful to explain.
Separating Understanding from Execution
One architectural decision I found important was keeping natural-language understanding separate from analytical execution.
Conceptually, the system can be viewed as several layers:
Natural-Language Query
↓
Query Understanding
↓
Analysis Planning
↓
Analytical Execution
↓
Evidence / Results
↓
Visualization
↓
Natural-Language Explanation
The language model is therefore not treated as the source of truth for visual analysis.
Its job is to understand the request and communicate the result. The specialized analytical pipeline is responsible for producing the underlying evidence.
A simplified representation of the routing logic looks like this:
query = "Show me the areas where vegetation decreased"
intent = understand_query(query)
if intent.type == "vegetation_change":
result = run_change_analysis(
image_a,
image_b
)
answer = explain_result(result)
The actual implementation can become considerably more complex, but the separation is important. It prevents the conversational component from becoming responsible for calculations and visual conclusions it cannot independently establish.
Making the Result Visual
Another important part of the system is that the result should not end as text.
If an analysis identifies regions of vegetation change, those regions need to be represented on the imagery or map. Conceptually:
result = analyze(images)
visual_layer = create_visualization(
image=images,
regions=result.regions
)
return {
"evidence": result,
"visualization": visual_layer
}
This gives the user two complementary forms of information.
The visualization answers:
“Where did this happen?”
The analytical result answers:
“What did the system measure?”
And the natural-language explanation answers:
“What does this result mean?”
Keeping these three responsibilities distinct makes the interaction easier to reason about.
Before and After: Turning a Query into an Analysis
Without this separation, a conversation might look like this:
User
Show me the areas where vegetation decreased between these images.
AI
Vegetation decreased in several areas between the two images.
That answer sounds reasonable, but it provides little evidence.
With the SatQuery approach, the interaction is intended to become:
User
Show me the areas where vegetation decreased between these images.
System
Identifies the request as a vegetation-change analysis.
Analysis pipeline
Processes the two images and identifies relevant regions.
Visualization
Highlights those regions on the satellite imagery.
System
Reports the resulting measurements and relevant confidence information.
AI
Explains what the analysis found in natural language.
The difference is subtle from the user's perspective, but significant from an engineering perspective. The second workflow has an explicit analytical stage between the question and the answer.
Adding Conversational Memory with Hindsight
Once the system can perform individual analyses, another problem appears: users naturally want to continue the conversation.
For example: User
Show me the areas where vegetation decreased between these images.
After receiving the result, the user might ask:
Now compare those regions with the previous analysis.
The second question depends on context from the first interaction.
I integrated Hindsight as the agent-memory layer for this part of SatQuery AI. Hindsight is designed to provide persistent memory for AI agents, allowing useful information from previous interactions to remain available instead of treating every interaction as completely isolated.
Hindsight on GitHub�
Hindsight Documentation�
Vectorize — What is Agent Memory?�
The distinction I found useful is that memory should preserve context, while the analytical pipeline should preserve evidence.
For example, memory can help establish that “those regions” refers to regions identified during an earlier vegetation analysis. It does not mean that memory itself becomes the source of the satellite-analysis result. Conceptually:
context = hindsight.retrieve(query)
analysis_request = understand_query(
query=query,
context=context
)
result = execute_analysis(analysis_request)
hindsight.remember(
query=query,
result_context=result
)
This makes the interaction conversational without collapsing the boundaries between memory, reasoning, and analysis.
The Architecture I Ended Up Thinking About
I think about SatQuery AI as six cooperating components rather than one large AI system.