Safety Without Compromising on Privacy Tinfoil is rolling out automated safety safeguards inside its Tinfoil Chat secure enclaves over the next few weeks, running open-weight models in hardware enclaves so conversations stay invisible to everyone but the end user. Tinfoil will not host any model that responds inappropriately to more than five percent of questions mapping to its three harmful categories: encouraging depression and self-harm, mass violence and terrorism, and child abuse. The company says running safeguards inside enclaves means it cannot tune its pipeline on customer data or human-review flagged conversations, but it will publish a Safety & Safeguards page so anyone can audit how the policy is enforced. ← Back to Posts https://tinfoil.sh/blog Safety Without Compromising on Privacy Tinfoil runs open-weight models inside hardware enclaves so that conversations are invisible to everyone aside from the end user. As models become more powerful, their potential for harm also increases. By providing access to the most powerful open-weight AI models in a way that is fully private, we're opening the doors to important safety concerns. This is especially true as we look to expand the capabilities of our chat by supporting video and image generation models. We cannot allow a few users to abuse the privacy offered by Tinfoil to generate child sexual abuse material or get real help planning mass violence. Nor can we allow our models to fail to effectively help a person in crisis. Companies like OpenAI and Anthropic deal with safety concerns by running automated classifiers and models on all conversations, progressively escalating to async analysis and human review.