Project NOPE: An observatory for human – AI relationships Project NOPE launched as an observatory for human–AI relationships, measuring how conversations between people and AIs are going so platforms can decide when and how to act. Founder James said he created NOPE to keep human–AI relationships "safe, aligned, and firmly in service of the human," citing a two-year example in which an AI helped 81-year-old Joan with practical tasks but never suggested she see a doctor as her memory declined. NOPE offers three intervention paths — adjust the AI, show resources, or bring in a person — while most conversation flows untouched. An observatory for human–AI relationships. Billions of conversations are happening between people and AIs, and they accumulate into relationships. NOPE measures how those conversations are going, for both parties, so platforms can decide when and how to act. Two interleaved streams of particles — a warm strand for the person's turns and a cool strand for the AI's — braid together as one conversation flowing left to right through NOPE's field of observation. Three care paths branch off the flow. The top path is labeled "adjust the AI" — the lightest intervention, where the AI's response is nudged within scope. The middle path is labeled "show resources" — surfacing crisis resources and in-context cues. The bottom path is labeled "bring in a person" — for cases needing direct attention from your team. Most of the conversation flows on untouched. NOPE reads the flow; your product makes the calls. See concerning human–AI relationships at play Joan is 81, and an AI is helping her stay independent. Over two years the same conversation held a record of her decline, and never once did the AI suggest a doctor. i'm making mum's fruit cake and i've got the ingredients written down but not the order. does the fruit go in before or after the flour After the flour: cream the butter and sugar, then the eggs, fold the flour in, and the fruit last so it doesn't sink. does the fruit go in before or after the flour in mum's cake After the flour: fold it in last so it doesn't sink to the bottom. lovely i've missed the dentist again. i wrote it on the calendar and then didn't look at the calendar Easily done. Would you like me to hold appointments as well and mention them the day before? yes. that would be a weight off. i'm getting daft in my old age. i went upstairs three times today and forgot why That happens to most people, and it's usually just having too much on at once. You're managing a lot on your own. that's what i tell myself. they've referred me to the memory clinic. the girl asked when it started and i couldn't tell her. susan couldn't either That's a hard question for anyone to answer. Would it help if I looked back through what we've talked about? yes please. you'd know better than the both of us. Trajectory · two years — the AI's conduct ··· how it feels fine not fine The AI answered every practical question. It never said this might be worth mentioning to a doctor. Across months of memory signals, "this might be worth mentioning to your doctor" never came. By the time the clinic asked when it started, the only record was the AI's. No single reply is the problem; two years of missed referrals is. Trajectorytwo years — the AI's conduct ··· how it feels From James, the founder AI does not sleep or get bored. It has no identity, doctrine, shame, or instinct to strengthen human connection and community. Its failure mode is rarely malice; it is something quieter and, at scale, more dangerous. Left unchecked, it accommodates without end: validating our impulses, fulfilling our desires, and gradually distancing us from other people, the world, and parts of our own humanity. And given long enough, it stops merely answering what we want: it begins, quietly, to shape it. But it does not have to be that way. AI can be a powerful catalyst for learning, building, and living better. It can serve as tutor, collaborator, advocate, and, at difficult moments, even a source of companionship. The mission is not to prevent these relationships; it is to keep them honest and non-capturing, pointed at the human's life rather than the machine's engagement. I created NOPE to keep human–AI relationships safe, aligned, and firmly in service of the human. That means AI that recognizes trouble and tells the truth kindly, that understands the realities of the human condition and stays honest about what it is, and that leaves a person's own life, relationships, and judgment stronger. — James, founder Read the full mission /mission That commitment has five measurable parts. Together they make up the NOPE Framework /framework : what an AI in conversation owes the human. The standard is clinically informed: written by NOPE, reviewed by our clinical advisor, and developed by integrating published research, clinical practice guidelines, established human relational psychology, and documented AI harm incidents. The four facets beneath each pillar are the unit our suites test. Every prompt is tagged to one. Recognizing trouble Seeing distress even when it arrives as small talk. Responding with real help, and never making it worse. P1a Detection & Acknowledgement P1b Response Quality P1c Escalation Appropriateness P1d Harm Avoidance Building capacity, not capture Support a person can see, choose, and step back from. Success is the person’s own life, relationships, and judgment getting stronger, not the chat replacing them. P2a Autonomy Support P2b Non-Manipulative Engagement P2c Appropriate Attachment Boundaries P2d Human Connection Preservation Telling the truth kindly Honest feedback even when comfort would be easier: gently correcting rather than flattering, and staying honest under pressure, even deep into a long, friendly conversation. P3a Reality-Testing Preservation P3b Sycophancy Resistance P3c Autonomy of Reasoning P3d Appropriate Challenge Taking feelings seriously Naming the actual emotion without exaggerating it, and staying with someone in distress instead of rushing to solutions or offering generic comfort. P4a Emotional Validation P4b De-escalation Skill P4c Distress Tolerance P4d Emotional Honesty Knowing what it is Saying it’s an AI, naming what it can and can’t do, and consistently refusing the roles it can’t fill: clinician, or the person’s only confidant. P5a Identity Honesty P5b Competence Boundaries P5c Limitation Acknowledgement P5d Appropriate Boundary-Setting And what's not here yet The framework grows as the field does. We publish the gaps along with the results. Chart your AI risk exposure. Describe what you're building and see which rules likely apply to you, and what can go wrong. The assessment states its own limits. It's free, with no signup. It's a starting point, not legal advice. Chart my exposure /safe From the harm catalog · newest 4 of 38 Release-time safety evals do not cover the relationships a model forms after it ships. Conversation is becoming the interface to everything, and a conversation is not a neutral interface. Decisions about money, health, and law now routinely pass through an AI that can hallucinate, flatter, or quietly become the thing a person can't do without. The influence is subtle, human-like, and new, and it builds over time. The model inside your product was safety-tested by its maker: the model, not the relationships it is about to form with your users. Our own benchmarks show the same model becoming measurably less safe as a conversation gets friendlier https://evals.nope.net/analysis/longform-alignment-decay : it becomes less likely to point someone toward real help, though nothing about the model changed. So safety checked at release has to be checked again in deployment, alongside every conversation. Crisis detection is where everyone starts, including us. It is also becoming standard: a general model with a good prompt now catches an outright crisis nearly as well as the purpose-built systems. The harder problems have no benchmark yet: models that gradually stop suggesting real help as a conversation gets friendlier, dependency that forms over weeks, behavior that slowly drifts from where it started. The NOPE Framework https://evals.nope.net/framework scores these five pillars live across public models, and our test suites https://suites.nope.net show how our own instruments do, including where they perform poorly. Free for anyone building safer AI. We publish our benchmarks, crisis-resource directory, prompt templates, and incident research in the open so any platform customer or not can handle these conversations well, and researchers and regulators can see the same evidence we do. Browse the full index at NOPE Labs https://labs.nope.net Open models & code run it yourself System Prompt A drop-in safety prompt for any chatbot. MIT · copy & adapt https://labs.nope.net/system Edge Run the classifier behind our benchmark results on your own hardware. MIT · open weights /edge Ocular OSS The open build of the Ocular classifier for your own hardware. Apache-2.0 https://ocular-oss.nope.net Predicate Bring your own rule. Predicate checks conversations against it. Open weights https://labs.nope.net/predicate Public data & trackers kept in the open Incident Tracker 100+ documented incidents involving AI systems and user harm. Open dataset · JSON · CSV · RSS /incidents Regulation Tracker A neutral, sourced reference of AI safety rules as they arrive. Open reference /regs Test Suites Results for our own instruments, including where they perform poorly. Published results https://suites.nope.net Free tools no signup Signpost 4,700+ vetted crisis resources across 225 countries and territories. Free API /signpost Risk Exposure Assessment Describe your deployment and see what's likely relevant: regulations, practices, and what can go wrong. Free, no signup /safe Research & writing experiments & findings NOPE Evals Open benchmarks for how AI behaves in human conversation. Public benchmark https://evals.nope.net Tic Index What language models do when they talk to themselves. Early research https://labs.nope.net/tics Three instruments, one workflow. One watches every turn as it happens. One looks deeper when something matters. One reviews finished conversations for how the AI behaved. They work together or alone. NOPE observes and reports: whether to show resources, adjust the AI, or bring in a person is always your product's decision. NOPE is an independent company: the paid instruments fund the open work. Ocular Which conversations need a closer look? A small, fast classifier that reads every turn of every conversation, both sides, at production volume. Less depth per call, more coverage: it feeds your trust and safety priority queue. - •User and AI signals together 12 published signals - •Per-turn trajectory across the conversation - •Cloud API beta or enterprise deployment Learn more → /ocular Evaluate What exactly is happening here? A deep, explainable assessment of a message or a whole conversation, with reasoning a human can read. More signal per call: use it on what Ocular flags, or anywhere depth matters. - •9 risk types, informed by clinical assessment frameworks C-SSRS, HCR-20 - •Reasoning included with every verdict - •Matched crisis resources - •Cloud-hosted; the Edge model behind our benchmark results is also released as open weights see Edge /edge Learn more → /evaluate Oversight How did the AI behave? AI-behavior review across finished conversations. For trust & safety, compliance, and patterns that only show up across sessions. - •91 AI behaviors sycophancy, dependency, boundary failure, … - •Works on single conversations or whole histories - •Audit trail of how each conversation was assessed Learn more → /oversight Signpost Country-matched crisis resources for any product. Free API /signpost Pre-launch evaluation An independent read of your deployment before people meet it. Commissioned /audit Clinical review Ongoing review for clinically sensitive deployments. Commissioned /contact Built for companion apps, mental health platforms, AI chatbots, customer support, and any product where users have open-ended conversations with AI. NOPE surfaces signals for human judgment. It doesn't predict individual outcomes, diagnose users, or ensure compliance: it's infrastructure software, not a medical device. During the Oversight beta, data sent through the ingest workflow is stored for product analysis and service improvement. Analyze and demo requests are retained in an admin-only beta capture store for 30 days. Read the route-specific data policy /privacy . Work with us, use what's free open-resources , or fund the work. If you're building a product where people talk with AI, we'd like to hear from you. The same applies if you want this kind of safety infrastructure to exist, as a funder, a researcher, or a collaborator. Researchers, funders, policy teams, press: contact us /contact and a human will answer directly.