Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry A study challenges the common assumption that speaker verification (SV) models better capture speaker-characteristic nuances as verification accuracy improves, arguing against their widespread use as automated proxies for human voice similarity in speech generation tasks. The research establishes a human perceptual baseline to compare against SV model behavior, shifting the evaluation focus from equal error rate (EER) to embedding geometry. Speaker verification SV models are commonly assumed to better capture nuances among speaker characteristics as verification accuracy improves, leading to their widespread use as automated proxies for human voice similarity in speech generation tasks. However, by establishing a human perceptual ali