cd /news/machine-learning/chips-character-histograms-and-posit… · home topics machine-learning article
[ARTICLE · art-76360] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

CHiPS: Character Histograms and Positional Signals for Lightweight Authorship Attribution in Romanian Texts

Researchers propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts that uses character histograms and positional signals without tokenization or pretrained language models. On a locked grouped ROST split of 400 files from 10 authors, CHiPS-F achieves 0.9310 accuracy and 0.9341 macro-F1, while a matched character n-gram SVM reaches 1.0000 accuracy, indicating the method's focus is on transparent, restricted evidence under strict leakage control.

read1 min views1 publishedJul 28, 2026

arXiv:2607.22884v1 Announce Type: new Abstract: We propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts. All reported experiments are closed-set: the true author is one of the candidate authors in the training data. CHiPS studies two complementary fingerprints of writing style: CH-SVM, a character-histogram classifier based on one-character marginal distributions, and FFT12-LR, a positional-signal classifier that represents selected characters and punctuation classes as impulse trains (binary indicator sequences over character positions) and extracts Fourier/Welch spectral descriptors. We also report CHiPS-F, a leakage-safe decision-level fusion variant, and an optional top-5 listwise reranker trained only on out-of-fold predictions. The method requires no tokenization, syntactic analysis, pretrained language model, or transformer fine-tuning, and it avoids character $n$-gram features with $n \geq 2$ in the histogram component. On a locked grouped ROST split comprising 400 files from 392 source-text groups, written by 10 authors, with source-text-level evaluation and grouped five-fold model selection, CHiPS-F reaches 0.9310 accuracy and 0.9341 macro-F1. A matched but unrestricted character 2--5-gram TF--IDF SVM comparator reaches 1.0000 accuracy and macro-F1 on the same held-out groups, so the contribution is not a claim of best possible classification accuracy. Instead, the experiments ask how far restricted, transparent character evidence can go under strict leakage control. On ROSTories-cleaned, a secondary ROST-overlapping corpus comprising 1,248 files from 1,240 source-text groups, written by 19 authors, the same protocol gives 0.8919 accuracy and 0.8708 macro-F1 for CHiPS-R.

── more in #machine-learning 4 stories · sorted by recency
── more on @chips 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/chips-character-hist…] indexed:0 read:1min 2026-07-28 ·