{"slug": "evaluating-llm-generated-rules-for-heart-disease-prediction", "title": "Evaluating LLM-Generated Rules for Heart Disease Prediction", "summary": "A study comparing traditional machine learning models against LLM-generated rule-based systems for heart disease prediction on the UCI Heart Disease dataset found that traditional models consistently outperformed the LLM-generated rules, with Random Forest achieving the best overall performance at 90.2% accuracy, 0.829 precision, 1.0 recall, and a 0.906 F1-score. The LLM-generated rule models trailed, with Claude Sonnet 4.6 reaching 80.3% accuracy (F1-score 0.833) and GPT-4o obtaining 70.5% accuracy (F1-score 0.690), though the LLM rules offered interpretable IF-THEN diagnostic logic that enhances explainability and transparency in clinical decision-making. The study, posted as arXiv:2609.13192v1, highlights the trade-off between predictive performance and interpretability in medical artificial intelligence systems, with the full implementation publicly available on GitHub.", "body_md": "arXiv:2609.13192v1 Announce Type: new \nAbstract: This study compares traditional machine learning models and Large Language Model (LLM)-generated rule-based systems for heart disease prediction using the UCI Heart Disease dataset. Several classifiers, including Logistic Regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Naive Bayes, Decision Tree, and Random Forest, were evaluated alongside rule-based systems generated using GPT-4o and Claude Sonnet 4.6. Model performance was assessed using accuracy, precision, recall, and F1-score metrics. Experimental results show that traditional machine learning models consistently outperform LLM-generated rule-based systems in predictive performance. Random Forest achieved the best overall performance with 90.2% accuracy, a precision of 0.829, perfect recall of 1.0, and an F1-score of 0.906. Naive Bayes followed closely with 88.5% accuracy and an F1-score of 0.881. In contrast, the LLM-generated rule models achieved lower performance, with Claude Sonnet 4.6 reaching 80.3% accuracy (F1-score: 0.833) and GPT-4o obtaining 70.5% accuracy (F1-score: 0.690). Despite the performance gap, the LLM-generated rules provide interpretable IF-THEN diagnostic logic that enhances explainability and transparency in clinical decision-making. These findings highlight the trade-off between predictive performance and interpretability in medical artificial intelligence systems. The complete implementation of all experiments, including machine learning models and LLM-derived rule classifiers, is publicly available in the GitHub repository at https://github.com/FeisalAlaswad/LLM-Rule-ML-Heart-Disease-Prediction .", "url": "https://wpnews.pro/news/evaluating-llm-generated-rules-for-heart-disease-prediction", "canonical_source": "https://arxiv.org/abs/2609.13192", "published_at": "2026-09-15 04:00:00+00:00", "updated_at": "2026-09-15 04:30:10.197474+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "ai-ethics"], "entities": ["UCI Heart Disease dataset", "GPT-4o", "Claude Sonnet 4.6", "Random Forest", "Naive Bayes", "Logistic Regression", "Support Vector Machine", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/evaluating-llm-generated-rules-for-heart-disease-prediction", "markdown": "https://wpnews.pro/news/evaluating-llm-generated-rules-for-heart-disease-prediction.md", "text": "https://wpnews.pro/news/evaluating-llm-generated-rules-for-heart-disease-prediction.txt", "jsonld": "https://wpnews.pro/news/evaluating-llm-generated-rules-for-heart-disease-prediction.jsonld"}}