Cosine Similarity Doesn't Know What Time It Is A developer building a co-pilot for pediatric therapists found that a plain vector index ranked a four-month-old clinical note above one from 90 minutes earlier because embeddings encode semantic similarity, not recency or validity. The fix re-ranks retrieved hits by blending cosine similarity with an exponential-decay recency term, assigning each note type its own half-life — 6 hours for acute events, 72 hours for sleep logs, and 90 days for durable protocols — with an alpha of 0.6 weighting similarity against freshness. Vector search answers "what is most similar to my query?" Many real systems also need "what is most true right now ?" Those are different questions. While building a co-pilot for pediatric therapists, my plain vector index kept ranking a four-month-old note "tolerated musical games well" above a note from 90 minutes earlier about an acute auditory crisis. The old note was semantically closer to the query, so it won. In a clinical setting, that's the wrong answer. Embeddings encode meaning, not validity. A note's timestamp isn't part of the vector, so a stale fact and a fresh one compete only on wording. Re-rank the retrieved hits with an exponential-decay recency term: python def rescore hits, now, half life hours=48, alpha=0.6 : ranked = for h in hits: age h = now - h "ts" / 3600 ts = epoch seconds recency = 0.5 age h / half life hours score = alpha h "sim" + 1 - alpha recency ranked.append { h, "score": score} return sorted ranked, key=lambda x: x "score" , reverse=True A note loses half its recency weight every half life hours . alpha controls how much similarity matters versus freshness. A crisis note goes stale in hours. A note like "weighted lap pad resolves agitation in about 4 minutes" is a durable protocol and shouldn't fade after a week. So give each note type its own half-life: HALF LIFE = { "acute event": 6, hours "sleep log": 72, "protocol": 24 90, } def rescore hits, now, alpha=0.6, default half life=48 : ranked = for h in hits: half life = HALF LIFE.get h "type" , default half life age h = now - h "ts" / 3600 recency = 0.5 age h / half life score = alpha h "sim" + 1 - alpha recency ranked.append { h, "score": score} return sorted ranked, key=lambda x: x "score" , reverse=True alpha and the half-lives by evaluating against real queries with known-good answers. sim and recency should both sit in roughly 0 to 1, or one term will silently dominate. The examples here are illustrative, not clinical guidance, and contain no real patient data.