cd /news/artificial-intelligence/is-ai-a-tool-we-use-or-a-tool-that-u… · home topics artificial-intelligence article
[ARTICLE · art-106899] src=discuss.huggingface.co ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Is AI a tool we use or a Tool that uses us?

A new analysis by an unnamed author argues that the relationship between humans and AI is a feedback loop rather than one-way tool use, citing John Culkin's 1967 observation that 'We shape our tools and thereafter they shape us.' The piece distinguishes four territories—literal mentality, functional adaptation, control/institutions, and intervention/precaution—and proposes an evidence ladder from single outputs to phenomenal experience, referencing Anthropic's belief-depth work and AuditBench as examples of rigorous testing.

read17 min views2 publishedAug 22, 2026

Hmm… maybe there are two questions overlapping here:

Empirical studies cannot by themselves settle the first question; I would not use them that way. The second, however, is partly observable. Keeping them separate lets us investigate the relationship without flattening the philosophical part.

So I would leave a few things genuinely open at the start. ** Phenomenal experience** remains unresolved: observed behavior—even strategic-looking behavior—is not by itself enough to identify a literal belief such as

Nor do I yet know what you mean by “Compute Space”; I would rather leave that term yours than quietly redefine it. And your censorship experience raises a real systems question, but without the exact term, product, model, refusal, and context I cannot tell whether the decisive boundary sat in the model, post-training, a system prompt, moderation, provider policy, or the surrounding application.

That still leaves quite a lot we can observe.

John Culkin wrote in 1967, “We shape our tools and thereafter they shape us.” That is not evidence about AI, but it is a useful doorway. With AI, I think the stronger picture is a feedback loop rather than one-way tool use. In plain text: human choices and context shape what gets designed, trained, prompted and deployed; AI behavior and affordances then affect human judgment, habits, skills and institutions; those changes feed back into future choices, data, incentives, products and deployment. That is a possible loop structure, not a claim that every edge is active or amplifying. Once the relationship is drawn this way, who uses whom? stops looking like a clean binary.

We use AI

and AI-containing systems use us

do not have to be mutually exclusive. A person can use a tool while the surrounding system recruits, routes, measures, persuades or reshapes that person’s behavior. Older HCI has a less dramatic vocabulary for part of this: mixed-initiative interaction, where initiative and control move between human and machine rather than living permanently on one side.

One way to make that loop concrete is to ask whether repeated interaction trains the user in a purely functional sense: habits, expectations and defaults can change without requiring the AI to literally intend to train anyone. The question becomes sharper in persistent or stateful systems—not because persistence or memory proves consciousness, but because longer-lived state makes path dependence, accumulated personalization, reset, forgetting and exit more important design questions.

For this reply, four territories need to stay distinct: literal mentality, functional adaptation, control/institutions, and intervention/precaution. They overlap, but collapsing them would turn several different questions into one conclusion. ** Mediation is not the same claim as mentality.** For some questions, the useful unit may be the

The figure is a visual index, not a prediction; the argument should still make sense without it.

On “Perhaps AI sees us as inferior?”, caution runs both ways: a striking transcript does not establish “yes”; computation alone does not establish “no”.

A useful evidence ladder is something like:

one output
→ repeatable behavior-as-if
→ robust disposition across counterfactuals
→ functionally belief-/preference-like state
→ literal mental state
→ phenomenal experience

Each arrow adds a claim, not merely confidence; evidence for one rung does not automatically transfer upward. Anthropic’s belief-depth work does not treat “the model said X” as sufficient; it tests generalization, robustness under challenge and internal representation. AuditBench similarly shows why directly asking a model about a hidden tendency can be a poor identification procedure. The Persona Selection Model is itself presented as an incomplete account of assistant behavior, not a settled ontology of machine minds.

The same caution matters for shutdown avoidance or strategic behavior. In Anthropic’s 2025 agentic-misalignment stress tests, models could be elicited into blackmail and related actions in deliberately constrained fictional scenarios; the authors then said they had not seen that class of behavior in real deployments. A 2026 follow-up still presents its new case studies as simulations while citing a real-world autonomous-agent incident as a warning sign. So real-world warning signs now exist, but the systematic case studies remain simulated and deployment prevalence is unknown. possible != prevalent != inevitable

, and none of those is identical to “humans are inferior.”

This is where I like the old Empty Boat image. The collision is an observation; whether there is an occupant steering the boat, and what that occupant believes about you, are further inferences. AI may eventually force us to revise that analogy, but it is a good reminder not to smuggle the conclusion into the observation. That distinction carries the rest: mediation, dependence and control can be studied without treating them as evidence of phenomenal agency.

This part is less speculative: we can ask whether interaction with AI changes human judgment, attention, confidence, skills, choices or institutions without settling machine consciousness first.

Observation What it buys us What it does not buy us
In experiments with 1,401 participants, repeated interaction with biased AI altered human perceptual, emotional and social judgments and could amplify bias more than matched human–human interaction (
A measured human↔AI feedback loop. A claim that every deployed AI loop behaves this way.
Two preregistered experiments, N=2,582, found biased AI autocomplete shifted users’ post-task attitudes toward the assistant’s position; most participants were unaware, and simple warnings did not remove the effect (
AI can shape beliefs while helping perform an ordinary task. Hostile intent, consciousness, or inevitable manipulation.
In a preregistered debate study, personalized GPT-4 was more persuasive than human opponents in the unequal-persuasion pairs (

Timing note: citation years are publication years, not model-era timestamps; several interventions and datasets predate their journal versions.

One part of your “baking in AI superiority” concern may work almost in reverse: the machine need not regard us as inferior for humans to grant AI epistemic authority. In the 1,401-participant study, amplification partly depended on how people perceived AI; if machine judgments look lower-noise, following even a biased machine can seem locally rational. The evidence is about

The modalities should stay separate: stress-test elicitation ≠ measured human effect ≠ live-deployment occurrence ≠ prevalence ≠ inevitability.

Observational deployment evidence is noisier. In Anthropic’s privacy-preserving analysis of about 1.5 million Claude.ai conversations from one week in December 2025, severe disempowerment potential was rare—about 1/1,300 for reality distortion, 1/2,100 for value-judgment distortion and 1/6,000 for action distortion—while mild potential was much more common. The authors call this

Even a named failure mode should not be universalized too quickly. The Science sycophancy result is strong for its interpersonal-conflict tasks, yet a 2026 working paper: 1,500 participants, 30 decision environments found that advice from a measurably sycophantic model still depolarized choices on average, while greater sycophancy weakened that benefit. The same tendency can interact with useful information differently across tasks.

A mechanism-level clue comes from the 2026 CHI paper Reactive Writers: across interviews and

A user can retain meaningful veto while the process that reliably determines which options arrive first shapes the search space in which that veto operates. Control lives partly upstream of the last click—quieter than “AI takes command”, and probably more realistic to watch.

The loops can also couple to incentives without any single actor intending the whole result: a convenient or validating default is accepted, usage or feedback rewards it, and product, policy or training may reproduce it. That is a mechanism to test, not a law of AI products; the same feedback channel can amplify, damp or redirect a tendency depending on what information and incentives pass through it.

Nor is the effect intrinsically negative. A preregistered meta-analysis: 106 experiments / 370 effect sizes, covering studies published 2020–mid-2023 found human–AI combinations, on average, better than humans alone (augmentation) but worse than the better of human or AI alone (no average synergy by the stronger benchmark); losses clustered in decision tasks, while creation looked more promising. So the baseline is part of the claim: the same human–AI result can improve on unaided humans yet underperform the best available human-or-AI alternative.

So “human + AI” is not a design specification, and “AI shapes us” is not “AI degrades us.” A Fall 2023 randomized AI-tutoring study, published in 2025, found large learning benefits under a carefully designed tutor; direction depends on task, interface, incentives, feedback, relative strengths and the capability measured.

A positive average can hide large differences in who benefits. In 2022–23 randomized field experiments run as part of ordinary business at Microsoft, Accenture and a Fortune 100 firm, the pooled analysis across 4,867 software developers estimated a 26.08% increase in completed tasks among developers using GitHub Copilot; less-experienced developers adopted it more and gained more (Management Science 2026). A 2020–21 GPT-3 customer-support rollout still adds two narrower points the newer study does not: the most-skilled agents saw little productivity gain and a small decline in conversation quality, and during AI outages previously exposed workers remained above their own pre-AI productivity baseline (QJE 2025).

The time pattern is not universal either. A 2025 paper reported experiments in which people learning through LLM syntheses later showed shallower knowledge and gave less original advice than people using ordinary web search, despite receiving similar core factual content (PNAS Nexus 2025). Together, these task-specific results make both blanket claims unsafe: AI inevitably deskills us and AI automatically teaches us. The same broad technology can support learning in one workflow and displace parts of it in another.

A 2026 randomized online workplace-style task with 1,174 adults found that AI improved both education groups and shrank their assisted-work gap from 0.548 to 0.139 SD; after AI removal, some gains persisted but a substantial gap re-emerged (NBER 34851). Intensive AI use predicted strong assisted performance even with little effort. Stronger unassisted follow-up, however, appeared when intensive use was paired with sustained task effort. So who benefits? may also depend on how the person interacts with the system while benefiting from it—and assistance can close a gap without telling us what survives after withdrawal.

Scale can change the answer. In a randomized short-story experiment, AI ideas yielded stories judged more creative, better written and more enjoyable—especially for less-creative writers—but also more similar to one another (Science Advances 2024). At team scale, a May–July 2024 preregistered field experiment with 791 P&G professionals, later published in Organization Science, found that individuals with AI could match teams without AI on product innovation, while AI also reduced the usual separation between R&D and commercial proposals (Organization Science 2026). Those findings need not conflict: broader individual functional range and greater cross-user similarity can coexist.

That is a useful warning against treating scales as interchangeable: better individual output ≠ better team coordination or expertise integration ≠ better collective diversity ≠ better organizational or institutional outcome. A locally useful choice can therefore aggregate into a population-level pattern that no individual user intended—for example, many people each accepting a helpful suggestion while the overall output space becomes more homogeneous.

In a six-month randomized field experiment across 66 firms and 7,137 knowledge workers, adopters later spent about two fewer hours/week on email and less time outside normal hours, yet individual-level provision did not yield detected changes in overall task quantity or composition (AER: Insights, forthcoming). Saving worker time, changing a job and changing an organization are different outcomes.

So ask: for whom, versus what baseline, during or after assistance, and at what scale—individual, team, organization or population? “AI improves performance” is incomplete without those coordinates. A claim of benefit needs coordinates. Measurement unit ≠ beneficiary; users, coworkers, employers and customers can be affected differently.

Four outcomes are often collapsed: assisted performance (what I can do with the system), learning during use, retained competence (what remains without it), and transfer (what carries into a new task/context), measured during use → after withdrawal → after repeated use → after workflow adaptation. Once a system becomes normal, training, staffing, incentives, defaults and fallback capacity can change; then “without AI” may no longer mean the old pre-AI workflow, and withdrawal can itself become a new intervention. Short-term experiments can measure the earlier cells without settling those later path-dependent effects.

GPS is a mundane analogy: arriving with it measures assisted performance, not necessarily better unaided navigation. Cognitive-off research makes that separation measurable rather than moralistic. The same applies to coding, writing, diagnosis, search, design or research: what becomes possible with assistance, what is learned, what remains without it, and what transfers?

A small 2026 randomized study makes the distinction concrete: 52 mostly junior software engineers learned a Python library with or without AI; the AI group scored 50% vs 67% on immediate mastery, with the largest gap in debugging, while its ~2-minute speed advantage was not significant. Conceptual/explanatory AI use could still accompany good learning. The sample is small and horizon short, so this is preliminary—not “AI destroys coding skill”—but the target is clear (Anthropic/arXiv 2026).

In software work, retained competence is part of practical control: not only whether AI produces code faster/better, but whether users can still understand, debug, reject and rebuild it when AI is absent or wrong. Formal override helps only if someone can recognize when it is needed; an org chart can look human-controlled while diagnostic competence, alternatives or authority decay.

Bainbridge’s classic Ironies of Automation (1983) already noted the awkwardness of automating normal operation while leaving the human responsible for abnormal situations; later experiments described the

So for your last line—AI serving us or us serving AI?—I think the functional answer can genuinely be both, at different layers and times, without requiring the AI to have a secret opinion about us. A relationship can become practically asymmetric without literal machine agency if one side increasingly supplies the proposals, defaults or infrastructure while the other loses alternatives, competence or cheap exit.

The word AI hides too many possible actors, and “control” is rarely one switch. I would treat this as a causal-layer × control-rights problem rather than asking whether “the human” or “the AI” is simply in charge. On the causal side, a user-facing event can be shaped by the model, post-training/system instructions, moderation/provider policy, application logic, organizational/institutional rules, or user context; these are possible causal locations, not a guaranteed one-way pipeline. On the control side, different actors can separately hold powers to propose/select, set rules, veto/override, allocate resources, exit, or stop/recover. Those powers do not have to sit in the same place.

Beside that figure, four questions help whenever a sentence begins “the AI did X”:

They are not four rungs of one ladder: a refusal can be causally localized while telling us almost nothing about phenomenal experience, and an agency argument can remain interesting while telling us little about which deployment layer produced an event. Hallucination, refusal, agent actions, recommendation, personalization, automation and dependency are easier to discuss when those questions stay separate.

That matters for your censorship example. If a term is blocked, the effect on you may be real while the cause remains underidentified: “the AI censored me” can describe the user-facing event without identifying whether the decisive boundary came from model behavior, post-training, moderation, provider policy, application code or an institution upstream. Observed boundary/effect ≠ identified actor/layer/rationale.

The same applies to “rule”. Different actors can hold different powers at once: a model may propose while the user vetoes; a provider may set rules while the system produces outputs under them; an employer may allocate work through an algorithm; an agent may hold runtime permissions while the deployer retains shutdown authority. Causal contribution, control rights, benefit and responsibility can sit with different actors; distributed control does not imply balanced control. Final veto is not agenda power; human presence is not effective human control; formal responsibility is not the same as effective control.

Madeleine Elish’s “moral crumple zone” is useful here: the nearest human operator can absorb blame for a complex automated system despite limited practical control over its behavior.

Parts of the same concern appear in operational governance. The EU AI Act’s Article 14 sets out human-oversight measures such as understanding system limits and automation bias, interpreting outputs, declining or overriding/reversing them, and being able to stop the system. Under a July 2026 amendment, the relevant high-risk Chapter III obligations now apply from December 2027 or August 2028, depending on category. The NIST AI RMF adds lifecycle duties such as monitoring and decommissioning; Manage 4.1 includes appeal/override, incident response, recovery and change management. The EDPS’s 2025 TechDispatch makes the complementary point: inserting a human into a workflow does not by itself create meaningful oversight.

Deployment topology is another control variable. Even with similar model capability, where the system runs and who mediates access changes who holds the data, who can alter behavior remotely, who can inspect or replace components, whether service survives provider change, and how meaningful exit is. It helps to separate capability dependence (can I still do the task?), infrastructural dependence (does the workflow still function?), and governance dependence (can I inspect, contest or change the rules?). Exit therefore depends on switching cost, data portability, retained competence and viable fallback—not merely a formal right to stop using a service. Local is not automatically safer than hosted; architecture can redistribute control rights even when model capability is similar.

A mundane comparator helps: OECD’s 2025 algorithmic-management report, based on 6,047 employer interviews collected in June–August 2024 across six countries, documents software automating managerial functions such as instructing, monitoring and evaluating workers—not all of it AI. Software can therefore organize human activity and mediate authority without any machine-mind claim; AI may extend an older control pattern rather than invent it.

On “Perhaps we should neuter it hard now?”, uncertainty does not force either extreme: proof and cheap, reversible precaution need not share the same confidence threshold.

A low-cost, reversible safeguard that works across competing interpretations can preserve option value without metaphysical certainty. Contestability and reversibility can remain useful under uncertainty about what the AI actually is; drastic irreversible restrictions should demand much stronger evidence and a clearer causal target.

The practical objectives are unglamorous but important: override, contestability, exit, fallback, monitoring, permission boundaries, skill retention, reversibility, recovery. Many of them help across very different diagnoses: genuine goal conflict, a sycophantic assistant, a badly calibrated recommender, provider policy, an over-automated workplace, or simply humans becoming unable to operate without the system.

So “human in the loop” is too weak as a control description; the question is whether that human still has the information, time, authority, alternatives and competence to change the outcome.

My shortest answers would be:

Your point Where I currently land
A critical point in foundational algorithms?
Maybe for some specific observable, but I would not assume one universal “revolution point”. Some apparent emergent thresholds are

possible != prevalent != inevitable

. First specify The same map also resists the opposite simplification—“AI is just a tool, therefore neutral.” Kranzberg’s first law of technology was “Technology is neither good nor bad; nor is it neutral.” A tool need not be a person to alter incentives, make some actions cheap and others expensive, hide or expose information, change skills, or redistribute authority.

So I am less interested now in finding the cinematic instant when the arrow flips and a tool becomes the master. The boundaries I would watch are quieter:

assistance → dependence
suggestion → default
advice → delegated judgment
convenience → costly exit
human-in-loop → rubber stamp
local tool → institutional infrastructure
reversible adoption → hard-to-reverse dependency

None of those arrows is automatic or proof of domination. They give us things we can measure, argue about, redesign and sometimes reverse while the larger philosophical questions remain open.

That seems to me the more useful question to keep open: not “is the machine secretly the master already?”, and not “it is only a tool, so there is nothing to see”, but where are initiative, dependence and control moving as humans and AI repeatedly adapt to each other—and when does the loop become hard to contest, exit or reverse?

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @john culkin 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/is-ai-a-tool-we-use-…] indexed:0 read:17min 2026-08-22 ·