Hey everyone, I 'm new here. Hope you guys don’t mind answering basic questions.
I’ve been exploring NSFW AI detection lately, and it’s been a pretty fascinating rabbit hole. Tools like NSFWJS are great for quick setups, and the CLIP-based NSFW Detector is super impressive with how it uses embeddings to classify content.
Recently, I came across this site called soulfun.ai (which is all about creative AI stuff including ai generated photos and videos), and it got me thinking: how can I fine-tune these models for more niche or specific datasets?
I’ve been playing around with a basic CLIP setup, and here’s a quick snippet of what I’ve tried so far:
from transformers import CLIPProcessor, CLIPModel
import torch
model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")
inputs = processor(text=["NSFW", "SFW"], images=image, return_tensors="pt", padding=True)
outputs = model(**inputs)
logits_per_image = outputs.logits_per_image # Scores for image-text similarity
probs = logits_per_image.softmax(dim=1) # Probabilities for each class
is_nsfw = probs[0][0] > 0.5
Please let know, for those of you who’ve fine-tuned a CLIP-based model for NSFW (or even something similar):
- What kind of datasets worked best for you?
- Did you use any specific tricks during training to improve accuracy?
- Any tips for keeping the model fast and lightweight during inference?
Would love to hear what’s worked for you! Thanks in advance for any advice.
I found a document that describes some of the parameters used during the tuning process, even though it is a ViT model rather than a CLIP model.
The basic flow and libraries used are the same even when tuning a CLIP model. It’s just a different model class.
Thanks, I’ll take a deep look!
The ViT model used a batch size of 16 and a learning rate of 5e-5. Do you think these parameters would be a good starting point for a CLIP model as well, or would adjustments be needed due to differences in model architecture? Anyway, thanks a lot!
I don’t have much experience training models, so I don’t really know!
However, since the image processing part of CLIP is ViT, I think it’s probably fine.
Well, I think that the optimal values are something that you have to try and find out, so I think it’s more reliable to adjust them while actually training.
hahaha, I should stop being lazy and try it for myself, thanks. Well, training models and optimizing them is really like alchemy, I guess.
What other dataset do you use for such NSFW content detection apart from NSFWJS? Thanks
One thing I’d focus on besides fine-tuning is the quality and diversity of the dataset.
AI-generated images can look very different depending on the model, generation settings, editing, and compression, so a model trained on a narrow dataset can struggle with images it hasn’t seen before.
I’d also test different confidence thresholds rather than relying on a fixed 0.5 cutoff. Check false positives and false negatives separately, especially for borderline images.
For practical testing, can include both real and AI-generated images, then add edited, resized, and compressed versions to see how much the detection accuracy changes. That usually gives a better idea of how well the model will perform outside the training dataset.