arXiv:2610.03531v1 Announce Type: new Abstract: Authorship Attribution (AA) requires capturing fine-grained stylistic characteristics, making it particularly challenging in zero-shot (ZS) settings where no task-specific supervision is available. In this work, we investigate the effect of author representations on ZS AA by evaluating a label-only prompting baseline together with three author representation strategies: representative writing samples, LLM-generated descriptions, and style embeddings (LISA). The first three approaches perform attribution using LLM prompting, while the embedding-based approach uses style embeddings with cosine similarity. We investigate the influence of prompt design and propose a two-stage embedding-based attribution framework that combines candidate space reduction with embedding-dimension selection. The results show that label-only ZS AA is ineffective, while incorporating author-specific representations consistently improves attribution performance. Among the evaluated approaches, the proposed two-stage LISA framework achieves the strongest overall performance, whereas LLM-generated style descriptions provide a substantially more compact representation of author style at the cost of some attribution performance. These findings demonstrate the importance of author representation in ZS AA, while indicating that current open-source LLMs remain insufficient for robust attribution without more effective representation learning.
Author Representation Strategies for Zero-Shot Authorship Attribution: A Comparative Study of LLM-Based and Embedding-Based Approaches
A comparative study of zero-shot authorship attribution (ZS AA) found that the two-stage LISA framework, which combines candidate space reduction with embedding-dimension selection, achieved the strongest overall performance among the evaluated author representation strategies. The study, posted as arXiv:2610.03531v1, evaluated a label-only prompting baseline alongside representative writing samples, LLM-generated descriptions, and style embeddings (LISA), finding label-only ZS AA ineffective while author-specific representations consistently improved attribution. LLM-generated style descriptions offered a substantially more compact representation of author style at the cost of some attribution performance, and the authors concluded current open-source LLMs remain insufficient for robust attribution without more effective representation learning.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.