17:45
2026-08-10
promptcube3.com
artificial-intelligence
Can we stop just randomly mixing safety data into LLM
A new approach called DataRx treats LLM safety as a missingness problem by analyzing hidden representations to sample only the most needed safety data, reducing attack success rates from 59.23% to 13.โฆ