cd /news/machine-learning/ensemble-of-unsupervised-deep-learni… · home topics machine-learning article
[ARTICLE · art-85551] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Ensemble of Unsupervised Deep Learning for Clustering Imbalanced Tabular Data

A new arXiv preprint (2608.00346v1) reports that ensemble methods combining unsupervised deep clustering algorithms outperform individual methods on imbalanced tabular data, achieving higher accuracy, normalized mutual information, and adjusted Rand index across 16 binary datasets. The researchers propose two ensemble approaches—one aggregating assignments across embedding dimensions and another using majority voting—and suggest deep clustering as a robust alternative to supervised classification under class imbalance.

read1 min views1 publishedAug 4, 2026

arXiv:2608.00346v1 Announce Type: new Abstract: Data imbalance poses a major challenge in supervised classification, where the majority-class bias contributes to false negatives and overestimates classification accuracy. Unsupervised deep clustering can be immune to class imbalance because representation learning for clustering is performed without class labels. Deep clustering has been proposed for images, languages, and graphs, while its application to tabular data has only emerged recently. This paper is among the first to examine the performance of state-of-the-art deep clustering methods under varying levels of data imbalance. We introduce two novel cluster ensemble approaches: one aggregates deep clustering assignments across different embedding dimensions, and the other applies majority voting to the best-performing clustering algorithms. Experiments on 16 binary tabular datasets with varying and artificially induced levels of imbalance reveal distinct strengths of different deep clustering methods. On average, our ensemble methods outperform individual clustering methods in ACC, NMI, and ARI scores, offering greater resilience to data imbalance when identifying ground-truth classes without supervision. Therefore, in an imbalanced data scenario, deep clustering can serve as a strong alternative to supervised classification.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ensemble-of-unsuperv…] indexed:0 read:1min 2026-08-04 ·