Your LLM red-team report should be reproducible, not a screenshot
A developer has released a red-teaming kit that makes LLM safety evaluations reproducible by pinning a versioned probe corpus, raw prompts and replies, and a sha256 hash of the corpus, scoring results 0-100 for use as a …