A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models A study evaluating five large language models as zero-shot annotators of English song lyrics found that LLM-based measurement reliability varies by social construct, with self-esteem showing the strongest repeated-measurement reliability and seeking recognition the least stable. The research, posted on arXiv (arXiv:2609.04428v1), used repeated annotations of a large lyric corpus to assess consistency across runs, convergence across models, and transferability to supervised classification, concluding that stability and convergence should be reported before LLM annotations are used as scalable measurements in cultural analytics. arXiv:2609.04428v1 Announce Type: new Abstract: Large language models LLMs are increasingly used to annotate cultural texts at scales that are impractical for human coders. However, before their outputs are treated as measurements of latent social constructs, it is necessary to establish whether those measurements are reliable. This study evaluates five LLMs as zero-shot annotators of four social constructs expressed in English song lyrics: self-esteem, self-control, seeking belonging, and seeking recognition. Using repeated annotations of a large lyric corpus, we examine three properties of LLM-based measurement: consistency across repeated runs, convergence across models, and transferability of consensus labels to supervised classification. The findings show that LLM-based measurement is not uniformly reliable across constructs. Self-esteem exhibits the strongest repeated-measurement reliability across models, while seeking recognition is generally less stable; self-control and seeking belonging show intermediate but model-dependent reliability. Downstream classification further indicates that consensus LLM labels contain learnable signal, although transferability does not itself establish construct validity. Repeated-measurement stability and cross-model convergence should therefore be reported before LLM annotations are treated as scalable measurements in cultural analytics.