Vector search answers "what is most similar to my query?" Many real systems also need "what is most true right now?" Those are different questions.
While building a co-pilot for pediatric therapists, my plain vector index kept ranking a four-month-old note ("tolerated musical games well") above a note from 90 minutes earlier about an acute auditory crisis. The old note was semantically closer to the query, so it won. In a clinical setting, that's the wrong answer.
Embeddings encode meaning, not validity. A note's timestamp isn't part of the vector, so a stale fact and a fresh one compete only on wording.
Re-rank the retrieved hits with an exponential-decay recency term:
def rescore(hits, now, half_life_hours=48, alpha=0.6):
ranked = []
for h in hits:
age_h = (now - h["ts"]) / 3600 # ts = epoch seconds
recency = 0.5 ** (age_h / half_life_hours)
score = alpha * h["sim"] + (1 - alpha) * recency
ranked.append({**h, "score": score})
return sorted(ranked, key=lambda x: x["score"], reverse=True)
A note loses half its recency weight every half_life_hours. alpha controls how much similarity matters versus freshness.
A crisis note goes stale in hours. A note like "weighted lap pad resolves agitation in about 4 minutes" is a durable protocol and shouldn't fade after a week. So give each note type its own half-life:
HALF_LIFE = {
"acute_event": 6, # hours
"sleep_log": 72,
"protocol": 24 * 90,
}
def rescore(hits, now, alpha=0.6, default_half_life=48):
ranked = []
for h in hits:
half_life = HALF_LIFE.get(h["type"], default_half_life)
age_h = (now - h["ts"]) / 3600
recency = 0.5 ** (age_h / half_life)
score = alpha * h["sim"] + (1 - alpha) * recency
ranked.append({**h, "score": score})
return sorted(ranked, key=lambda x: x["score"], reverse=True)
alpha and the half-lives by evaluating against real queries with known-good answers.sim and recency should both sit in roughly 0 to 1, or one term will silently dominate.
(The examples here are illustrative, not clinical guidance, and contain no real patient data.)