Managing prompts for Large Language Models (LLMs) often feels like a dark art. A seemingly minor tweak to a system prompt can inadvertently cause unintended behavioral shifts or output formatting failures—a phenomenon known as semantic drift.
To address this, we need a way to statically analyze and quantify the impact of prompt changes before deploying them, without relying on slow and expensive LLM-in-the-loop evaluations for every minor iteration.
Below is the implementation of Agentic-Local-Prompt-Semantic-Diffuser, a lightweight, zero-dependency Python CLI tool designed to evaluate the semantic impact of prompt modifications in under 10 seconds. It utilizes term frequency vectorization and cosine similarity to measure semantic drift, alongside structural static analysis to detect formatting risks.
Save the following script as prompt_diffuser.py.
"""
Agentic-Local-Prompt-Semantic-Diffuser
A one-shot CLI tool to automatically evaluate the semantic impact of prompt changes for local LLMs in under 10 seconds.
[Execution Example]
python prompt_diffuser.py \
--old-prompt "You are a helpful assistant that outputs JSON." \
--new-prompt "You are a strict assistant that outputs strict JSON format." \
--test-inputs "Hello" "What is the weather?"
"""
import sys
import json
import argparse
import math
from typing import List, Dict, Any
def tokenize(text: str) -> List[str]:
"""Tokenizes the input text into a list of lowercase words."""
return text.lower().split()
def get_vector(text: str, vocabulary: List[str]) -> List[float]:
"""Generates a term frequency vector for the given text based on the vocabulary."""
tokens = tokenize(text)
return [float(tokens.count(word)) for word in vocabulary]
def cosine_similarity(v1: List[float], v2: List[float]) -> float:
"""Calculates the cosine similarity between two vectors."""
dot_product = sum(a * b for a, b in zip(v1, v2))
norm1 = math.sqrt(sum(a * a for a in v1))
norm2 = math.sqrt(sum(a * a for a in v2))
if norm1 == 0.0 or norm2 == 0.0:
return 0.0
return dot_product / (norm1 * norm2)
def analyze_structure(text: str) -> Dict[str, Any]:
"""Performs static analysis on the prompt to extract structural metadata."""
return {
"length": len(text),
"has_json_hint": "json" in text.lower(),
"has_markdown": "`" in text or "#" in text,
"line_count": text.count("\n") + 1
}
def main():
parser = argparse.ArgumentParser(description="Agentic-Local-Prompt-Semantic-Diffuser")
parser.add_argument("--old-prompt", required=False, help="Path to old prompt or prompt string")
parser.add_argument("--new-prompt", required=False, help="Path to new prompt or prompt string")
parser.add_argument("--test-inputs", required=False, nargs="+", help="Representative test inputs")
args = parser.parse_args()
old_p = args.old_prompt if args.old_prompt else "You are a helpful assistant that outputs JSON."
new_p = args.new_prompt if args.new_prompt else "You are a strict assistant that outputs strict JSON format."
inputs = args.test_inputs if args.test_inputs else ["Hello", "What is the weather?"]
vocab = list(set(tokenize(old_p) + tokenize(new_p)))
for inp in inputs:
vocab = list(set(vocab + tokenize(inp)))
old_vec = get_vector(old_p, vocab)
new_vec = get_vector(new_p, vocab)
prompt_similarity = cosine_similarity(old_vec, new_vec)
semantic_drift = 1.0 - prompt_similarity
old_struct = analyze_structure(old_p)
new_struct = analyze_structure(new_p)
structural_risk = "LOW"
if old_struct["has_json_hint"] != new_struct["has_json_hint"]:
structural_risk = "HIGH"
elif abs(old_struct["length"] - new_struct["length"]) > 200:
structural_risk = "MEDIUM"
test_evaluations = []
for inp in inputs:
inp_vec = get_vector(inp, vocab)
sim_old = cosine_similarity(old_vec, inp_vec)
sim_new = cosine_similarity(new_vec, inp_vec)
test_evaluations.append({
"input": inp,
"old_alignment": round(sim_old, 4),
"new_alignment": round(sim_new, 4),
"drift_delta": round(sim_new - sim_old, 4)
})
report = {
"status": "SUCCESS",
"semantic_drift_score": round(semantic_drift, 4),
"structural_risk": structural_risk,
"details": {
"prompt_similarity": round(prompt_similarity, 4),
"old_structure": old_struct,
"new_structure": new_struct,
"test_evaluations": test_evaluations
}
}
print(json.dumps(report, ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
If executed without any arguments, the script will run an evaluation using the default sample prompts and test inputs, acting as a quick sanity check.
python prompt_diffuser.py
The tool outputs a structured JSON report. The semantic_drift_score provides a normalized delta between the two prompts, while the structural_risk flag alerts you to potentially breaking changes (such as accidentally removing a JSON formatting instruction).
{
"status": "SUCCESS",
"semantic_drift_score": 0.4377,
"structural_risk": "LOW",
"details": {
"prompt_similarity": 0.5623,
"old_structure": {
"length": 54,
"has_json_hint": true,
"has_markdown": false,
"line_count": 1
},
"new_structure": {
"length": 68,
"has_json_hint": true,
"has_markdown": false,
"line_count": 1
},
"test_evaluations": [
{
"input": "Hello",
"old_alignment": 0.0,
"new_alignment": 0.0,
"drift_delta": 0.0
},
{
"input": "What is the weather?",
"old_alignment": 0.0,
"new_alignment": 0.0,
"drift_delta": 0.0
}
]
}
}
To evaluate your own prompt iterations against representative user queries, pass them via command-line arguments:
python prompt_diffuser.py \
--old-prompt "Summarize the text." \
--new-prompt "Provide a detailed bullet-point summary in Japanese." \
--test-inputs "Machine learning is a subset of artificial intelligence."
Quantifying prompt drift statically provides a vital first line of defense before committing prompt changes to your application. While vector-based similarity cannot replace semantic evaluation by an LLM (such as LLM-as-a-Judge paradigms), it excels as a rapid, zero-cost CI/CD gatekeeper. By integrating this lightweight script into your workflow, you can proactively detect unintended regressions, missing format instructions, or excessive scope creep in your prompts without incurring API overhead.
If this engineering log saved your production server (and your sanity), consider supporting our architecture on GitHub Sponsors.