cd /news/artificial-intelligence/analysis-of-prompt-engineering-for-d… · home topics artificial-intelligence article
[ARTICLE · art-121165] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Analysis of Prompt Engineering for Drug Toxicity Prediction

A new analysis from arXiv (2609.03635v1) finds that natural variance in large language models (LLMs) outweighs prompt fine-tuning for drug toxicity prediction, despite clinical trials in the UK costing up to £1.3 million with an approximately 90% drug failure rate. The study, which prompted LLMs to identify chemical properties for toxicity prediction, showed substantial performance improvements when using chemoinformatic code instead of LLM-generated values.

read1 min views1 publishedSep 4, 2026

arXiv:2609.03635v1 Announce Type: new Abstract: Clinical trials in the UK can cost up to {\pounds}1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence has been increasingly explored to aid in the prediction of drug toxicity, with extensive use of large language models (LLMs). However, LLMs can show considerable variation when minor changes are made to prompts, which raises concerns about their sensitivity to prompt engineering. Prompt engineering is used to optimise a prompt given to an LLM to generate the desired output. This paper proposes a method to analyse prompt engineering for drug toxicity prediction. The aim of the paper is to investigate the importance of prompt phrasing for drug toxicity prediction. LLMs were prompted to identify chemical properties of significance when predicting drug toxicity. Prompts were constructed to investigate; job role, prompt structuring, and rule interpretation. LLMs were then used to generate datasets, using the identified features from initial prompting, which were then passed to machine learning algorithms. The experiments show that the natural variance which occurs in LLMs outweighs any fine-tuning of prompts. There were, however, substantial improvements in model performance when using chemoinformatic code to extract features instead of using LLM-generated values. The proposed analysis methodology is applicable to a wide range of prompt types across different areas of bioinformatics.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
wpnews · · #developer-tools
BoardUI
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/analysis-of-prompt-e…] indexed:0 read:1min 2026-09-04 ·