arXiv:2608.21464v1 Announce Type: new Abstract: We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization in a standard CNN classifier without architectural modification. Using synthetic images of colored geometric shapes, we encode classes as flat string labels (e.g., "red-circle") with no explicit attribute decomposition, and exclude certain color-shape combinations from training entirely. We apply two distortion methods derived from Jaccard string similarity between class names: mixed labels (soft target distributions encoding inter-class overlap) and expanded dataset (false training samples with structurally motivated incorrect labels). Both methods induce the ability to predict unseen class combinations, and act at different levels: mixed labels activate the classifier for unseen combinations by exploiting the CNN's natural embedding structure, while expanded training improves the embedding factorization itself. A control with random (unstructured) false labels confirms that the effect depends on the structure of the distortion, not on noise per se. These results suggest that structured complication of training signals can influence both the internal organization of learned representations and their compositional interpretation - a principle that may underlie the role of natural language in cognitive development.
Complexity Induction: Compositional Generalization via Structured Label Distortion
A new study on arXiv (arXiv:2608.21464v1) shows that structured distortion of training data, termed 'complexity induction', can induce compositional generalization in a standard CNN classifier without architectural changes. Using synthetic images of colored geometric shapes with flat string labels like 'red-circle', the researchers applied two distortion methods based on Jaccard string similarity—mixed labels and expanded dataset—to enable prediction of unseen color-shape combinations. A control with random false labels confirmed the effect depends on the structure of the distortion, suggesting that structured complication of training signals can influence both internal representations and their compositional interpretation.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.