cd /news/artificial-intelligence/02-what-does-general-in-agi-actually… · home topics artificial-intelligence article
[ARTICLE · art-125963] src=rainbowcity.substack.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

02 | What Does “General” in AGI Actually Mean?

In the second essay of his AGI series, developer KunYuan argues that the word "General" in Artificial General Intelligence is dangerously underspecified, with researchers, labs, and the public using it to mean anything from broad task coverage to autonomous planning and real-world collaboration. He contends that treating ever-longer capability lists as proof of generality confuses the debate rather than advancing it, and calls for philosophical conceptual engineering to separate the distinct levels bundled into the term.

by read23 min views4 publishedSep 10, 2026
02 | What Does “General” in AGI Actually Mean?
Image: Rainbowcity (auto-discovered)

Hi, I’m KunYuan. Let’s continue with the second essay in my AGI series. In the first essay, I raised what I believe is the first question we need to answer if we want to understand AGI: when we say “AGI has arrived,” who or what exactly is the subject of that claim? Are we talking about a standalone LLM? An Agent system formed by combining a large model with a Harness? Or an entire system that includes the model, tools, Humans, and organizational processes? If we have not even clarified the subject, then we do not yet know what exactly has arrived. How, then, can we judge whether it is AGI?

In this essay, I want to continue from there and pursue a second important concept inside AGI: General. Once again, we need to take a philosophical scalpel to the concept, clarify its meaning, and separate the different levels that have been mixed together. Only by opening it up carefully can we understand what it actually means. And only then can we begin to develop a shared language through which broader agreement about the definition of AGI might become possible.

We talk about Artificial General Intelligence, or AGI, every day, and we all know that the G in AGI stands for General. In Chinese, the word is usually rendered as “通用,” which is why AGI is commonly translated as “通用人工智能.” But what exactly does “General” mean? What does this word actually refer to? I suspect that most of us carry some implicit understanding of it, and for many people the first intuition is probably still capability breadth: a general intelligence is one that can do many different things.

So let us think about this together. If a model can write essays, program, conduct research, generate videos, use tools, and operate computers, does that make it General? If the number of things it can do keeps increasing and its capabilities keep becoming stronger, does that necessarily mean it is moving along a definite path toward AGI? This is the starting point for everything I want to examine in this essay.

Over the past few years, almost every major breakthrough in model capability has been followed by some version of the same judgment: the model has become more general, and it has moved another step closer to AGI. This sounds entirely natural because, in everyday language, “general” often means being useful for many purposes or being able to do many things. But in my view, once General becomes part of AGI, a concept that needs to be defined, studied, and verified, “being able to do more and more things” is no longer enough. It tells us that the capability list has grown longer, but it does not tell us what kinds of change this intelligence can remain effective across.

In my view, much of the confusion surrounding AGI today is hidden inside this still-unclear word, General. Different laboratories, researchers, and members of the public may all use the same word while actually referring to very different things. Some use “general” to mean being able to complete many tasks. Some mean being able to adapt to unfamiliar problems. Others mean being able to plan independently and use tools. Still others have already included autonomous judgment, long-term action, and real-world collaboration in the concept. On the surface, everyone appears to be debating whether a particular system is AGI, while in reality they may not even be talking about the same kind of General.

Even many claims that “AGI is almost here” do not necessarily make clear what General means in the claim itself. We often see a startling new capability and naturally conclude that the system has become “more intelligent,” “more General,” or “closer to AGI.” But if the concept itself remains vague, then stronger capabilities may actually make the debate around AGI more confused rather than less.

So before anything else, we need to open this word back up through philosophical conceptual engineering.

If an AI Has Ten Thousand Skills, Is It General? #

Let us begin by imagining an extreme system. It has ten thousand skills. It can write poetry, translate, program, create spreadsheets, analyze contracts, plan travel, diagnose equipment, and call ten thousand different tools. As long as a Human selects the correct task, provides the prescribed input, and presses the corresponding button, the system can perform extremely well. This system would obviously be powerful, and its capabilities would obviously be broad. But if those ten thousand skills correspond to ten thousand predefined paths, ten thousand separate Skills, with the boundaries, steps, judgment criteria, and termination conditions of every task already specified by Humans, then what exactly do we have: general intelligence, or an unimaginably large warehouse of capabilities? If a task arrives in a completely unfamiliar form, can the system still adapt? If a Human gives it only the goal without writing out the full path, can it organize the task by itself? If several goals conflict and the existing rules no longer provide a unique answer, can it continue to form judgment?

As I kept following these questions, I came to see that the number of capabilities can answer only “how much can it do?” It cannot answer “when the world changes, can it still do it?” What General is really concerned with is not how many capabilities an intelligence has accumulated in a static world, but what kind of relationship that intelligence has with changes in tasks, environments, paths, and rules.

So I arrived at a fundamental judgment about General:

The essence of General is not the number of capabilities, but the relationship between intelligence and change

.Every Claim of Generality Must Specify Three Things

If General describes a relationship between intelligence and change, then a complete claim of generality cannot simply say, “this system is highly general.” At minimum, it must answer three questions: who or what is being evaluated, what kind of change is being crossed, and what effectiveness requirements continue to hold after that change? The first is the object. Are we evaluating a model, an Agent, or an entire system that includes Humans and tools? This is the subject problem raised in the previous essay. The second is change. What exactly changed? Did the type of task change? Did the input distribution change? Was the task path left unspecified? Or did the existing rules cease to provide a unique answer? The third is effectiveness. After the change occurs, what counts as the system still being effective? Is producing an answer enough, or must it still complete the task? Is one occasional success enough, or must performance remain stable within the declared scope? How much time, data, tool use, and Human assistance are allowed? What kinds of failure would invalidate the generality claim?

Without any one of these three elements, object, change, and effectiveness, General slides back into a vague marketing term. A system may perform extremely well across a thousand tasks, but if we do not know which object actually possesses the capability, what change it has crossed, or where the boundaries of success and failure lie, then saying it is “more general” remains only an overall impression. It has not yet become a judgment that can truly be compared and verified.

Following this line of thought, I separate four different meanings that are often mixed together under General today: G1 Skill Generality, G2 Transfer Generality, G3 Open Generality, and G4 Subjecthood Generality.

These are not four product levels arranged from low to high, nor are they a strict capability ladder that has already been proven. They address four different kinds of change. G1 deals with changes in task type. G2 deals with changes in distribution, environment, and task form. G3 deals with situations in which task framing and workflows are incomplete. G4 deals with a normative gap, where external norms can no longer uniquely determine action.

One point must be clarified in advance: in my framework, G1 through G4 are four forms of generality, not four forms of AGI. Satisfying any one of them shows only that a system possesses the corresponding type of generality within a specified scope. It does not justify declaring that AGI has been established. Even if evidence supports all four forms separately, we still need to ask whether those capabilities belong to the same AI Actor that persists across time, whether they have been organized into a complete Capability Loop, and whether there is a Governance Loop that runs along that same path.

So let us keep taking the concept apart. Simply memorizing four names means very little. We need to see where, exactly, each of these four kinds of change occurs

.

G1: When the Type of Task Changes, Can It Still Do It? #

G1 addresses the most intuitive kind of change: the type of task changes. One moment we ask an AI to write an essay, the next to program, then to do mathematics, analyze an image, organize a spreadsheet, or build a webpage. These tasks require different kinds of knowledge and capability.

If the same AI can move between these different tasks and maintain effective performance within the scope we have declared, then what it demonstrates is G1, or Skill Generality. The real question behind G1 is: when the category of task changes, can the system continue to perform effectively? Much of what has astonished humanity about today’s large models comes precisely from this breadth of capability. They have moved from handling only text to understanding images, speech, and code; from merely generating answers to searching for information, calling tools, and operating computers. Tasks that previously required many different pieces of software, specialists, and workflows can increasingly be handled through the same model interface. This is a real and important advance in intelligence. There is no need to diminish the value of G1 simply because I want to emphasize G2, G3, and G4 later.

But G1 tells us only what the system can do across the multiple task categories already included in the declared scope. It does not automatically tell us whether those capabilities will still hold when the tasks appear in unfamiliar forms, nor whether the system can organize its own next steps when Humans no longer provide a complete path. A warehouse of skills can keep expanding, but the size of the warehouse and the ability of intelligence to remain effective across change are still two different questions.

G2: The Task May Be the Same, but the Conditions Have Changed #

G2 addresses a different kind of change. Here the task itself may not have changed. What changes is the form, condition, or environment in which the task appears.

Suppose a system has already learned how to solve a certain kind of problem. When the way the input is expressed changes, when the familiar surface form disappears, or even when the system encounters a form it has never directly seen before, can it still recognize the underlying structure and transfer its existing capability into the new situation?

At this point, I am no longer asking whether it can perform the task. I am asking:

When a familiar task becomes unfamiliar, can its existing capability still hold?

This is G2, Transfer Generality. It concerns changes in data distribution, environmental conditions, modes of expression, or task form, and asks whether an existing capability can survive that unfamiliarity and continue to work.

This distinction can easily be hidden by the sheer number of capabilities. If a system has already encountered a large number of different forms in advance and simply learned separate ways to handle each one, it may have broader G1 coverage without necessarily possessing stronger transfer capability. To evaluate G2, we need to specify exactly what is new about the new condition, how many new samples and interactions the system used, how much time and how many resources it required, and over what scope it successfully adapted. Some kinds of transfer may require no new samples at all, some may require a small number of examples, and others may require continued learning within the environment. But none of them can be collapsed into the vague statement, “it solved a new problem.”

So G1 and G2 are not simply higher and lower levels on a single ladder. A system may be able to perform many tasks yet be poor at handling unfamiliar change; another system may operate only within a limited domain but demonstrate extremely strong transfer within that domain. G1 asks, “what other things can you do?” G2 asks, “when the world becomes unfamiliar, can you still do what you already knew how to do?”

And even if an AI can both do many things and transfer extremely well, we still have not gone very deep into the open world. Up to this point, the task itself, and even the way the task is organized, may still have been defined for the AI by Humans

.

G3: When No Complete Path Exists, Can It Form Its Own Path? #

Now suppose I want an AI to help me organize a global AGI conference two months from now. But instead of assigning tasks one by one, such as writing invitations, analyzing contracts, and designing the schedule, I tell it only: help me organize this conference.

That goal does not come with a complete workflow, and no one can write out every step from today until the end of the conference in advance. Guests may decline invitations. The venue may suddenly cancel. The budget may change. Partners may introduce new conditions. A problem that has never appeared before may suddenly become central to the entire event. The AI has to understand the current state, identify information gaps, form intermediate tasks, determine their order, choose tools and collaborators, and continually revise its original plan in response to real-world feedback.

At this point, it is no longer simply selecting a skill from a task list, nor merely transferring the same skill to an unfamiliar input. It is continually forming and revising the task framework between the goal and reality. This is G3, Open Generality. G3 concerns what happens when workflows and task organization are incomplete. The question is: when Humans have not written out the entire road from goal to result in advance, can the system organize how to proceed by itself?

If G2 is like continuing to drive along an unfamiliar road, then G3 is like finding a route when no complete road exists, revising that route, and perhaps even rethinking how the destination should be reached. A system that reaches G3 has already begun to display strong Agent capability: it can plan, call tools, adjust its next step in response to results, and continue advancing a task over longer periods of time. But as I continued asking questions, a deeper boundary appeared:

Being able to organize how to do something is not the same as being able to judge what should be chosen.

Humans can set goals, rules, and priorities for AI, and then hand over the concrete organization of tasks. For example, we could stipulate that the lowest cost must always take priority, that any budget overrun must be rejected, that every option must be ranked according to a predetermined scorecard, and that any dispute not covered by the rules must be returned to a Human. In that case, the AI can still have very strong G3 because Humans have not written out every individual step. But what ultimately determines priority is still either the rules Humans provided in advance or the Human who takes over when a choice must finally be made.

So an open path is not the same as open judgment. A system may organize tasks by itself while still never truly confronting the moment when “the existing rules cannot decide for me.” And sooner or later, the real world will push any intelligence that acts continuously into exactly that position

G4: When the Rules Cannot Give a Unique Answer, Who Forms the Judgment? #

By the time we reach G4, the issue is no longer simply that the task has become more difficult. The change has occurred somewhere else.

Suppose an AI already possesses the relevant facts and understands the consequences of several possible directions. One option is safer but more expensive. Another is more efficient but requires taking greater risk. Different goals conflict with one another, and the existing rules do not specify a unique order of priority.

If facts are missing, the system can investigate further. If computation or resources are lacking, it can call more tools. But when the facts are already sufficient and the existing rules still cannot uniquely determine the action, adding more facts and compute cannot make the final choice on the system’s behalf. We call this position a normative gap. It may arise because rules are missing, because rules conflict, or because existing rules still cannot determine what should be done in the present situation. A normative gap does not mean the system knows nothing. It means that even when the relevant facts are known, external norms still do not provide a unique answer. What is truly missing is this: among several directions that are all actionable and all supported by reasons, which one should be chosen?

As long as AI continues to enter the open world over time, this kind of gap can never be completely eliminated in advance. Humans can keep improving rules and establishing boundaries, but they cannot foresee every future situation, nor can they write a unique answer in advance for every conflict of values. At this point, the question is no longer merely how to complete a task. It becomes: when the rules cannot decide, who forms the judgment? G4 is forced into view precisely here.

G4 is Subjecthood Generality. It addresses normative gaps in which external norms cannot uniquely determine action and asks: when external rules cannot provide a unique answer, can the system continue to form judgment, and to whom does that judgment actually belong?

If every time a normative gap appears, a Human ultimately chooses among the available options while the AI merely provides information, lists possible plans, or executes a decision already made by the Human, then the system may still possess excellent G1, G2, and G3. But the judgment within the normative gap cannot be directly attributed to the AI itself. An AI proposing several options does not mean the AI has adjudicated among them; filtering out obviously invalid options does not mean the judgment among the remaining options belongs to the AI; and the node that eventually issues the execution instruction is likewise not necessarily the subject that formed the judgment. Conversely, G4 does not require an AI Actor to make every judgment alone. Humans can provide goals and intentions, establish value and permission boundaries, and bring their own judgments into a process of joint deliberation. What we need to establish is whether the AI Actor genuinely participates in forming the judgment within the normative gap, and whether its contribution to the choice is identifiable, attributable, and actually plays a role in what is ultimately selected.

A Human retaining authorization over real-world action does not mean the AI failed to form its own judgment. If the AI has already formed a judgment and the Human is responsible for deciding whether that judgment qualifies to enter reality, then judgment and authorization occur at two different positions. This is also another foundation for the principle I proposed in the previous essay:

AGI may judge freely, but it may not act freely.

So when someone claims that an AI possesses G4, we still have to ask: what normative gap did it actually face? Why did the existing rules not already determine the answer? Who formed the judgment in this case? Did the AI merely generate an explanation that sounded reasonable, or did those reasons actually participate in the formation of the judgment at the time?

Following these questions downward eventually forces G4 back to the root question raised in the previous essay:

Who?

That is why we call G4 Subjecthood Generality. Here, “subjecthood” does not first mean claiming that AI already possesses consciousness, personhood, or legal identity. It points instead to a more basic structural fact: if we say that a judgment belongs to AI, then that judgment must be stably attributable to a subject that persists across time and can carry both the judgment and its consequences.

G4 is not a philosophical decoration that we arbitrarily attached after G1, G2, and G3. It is the question that inevitably appears when the open world pushes intelligence to the boundary of rules. And so, in the end, the problem of General returns once again to “who.”

The Four Forms of Generality Cannot Substitute for One Another #

At this point, the ontological boundaries of G1 through G4 become clear. G1 deals with changes in task type. G2 deals with changes in distribution, environment, and task form. G3 deals with incompleteness in task framing and workflows. G4 deals with situations in which norms can no longer provide a unique answer. G1 asks whether the system can do different things. G2 asks whether an existing capability still holds when conditions change. G3 asks whether the system can organize its own path when the road has not been written in advance. G4 asks who forms the judgment when the rules themselves cannot decide.

Because they address four different kinds of change, evidence for one form of generality cannot automatically serve as evidence for another. Strong results across many benchmarks may support G1 within a certain scope, but they do not automatically establish G2. Completing an unfamiliar task may support G2 within a corresponding scope, but it does not automatically show that the system can form an open task framework. Autonomous planning, tool use, and path revision may support G3, but they do not automatically show that the judgment within a normative gap was truly formed by the AI. And a single answer that appears highly opinionated does not automatically prove that the same AI can continue bearing those judgments across time.

The four forms of generality therefore carry different evidentiary obligations. G1 must provide evidence of capability coverage, stability, and failure boundaries for the same object across multiple kinds of tasks. G2 must provide evidence of transfer across unfamiliar distributions, environments, or task forms. G3 must provide evidence that the system itself forms and modifies tasks and paths when there is no complete predefined workflow. G4 must provide evidence that a genuine normative gap exists and show how the AI and the Human each participated in the formation of judgment. We cannot use the number of tasks to prove transfer, cannot use final success to prove that the path was formed by AI, and cannot use an answer that sounds highly opinionated to prove that the judgment truly belongs to AI.

Nor do these four forms of generality require every system to advance step by step from G1 to G2 to G3 to G4. A system may possess very strong transfer capability within a limited domain, or may be capable of attributable adjudication within a local normative gap while lacking broad skill coverage. The establishment of local G4 certainly does not mean that the system has already become an AGI across all domains.

What we are really trying to establish is not a simple ranking table, but a language that does not permit conceptual substitution. Anyone can propose their own definition of AGI, and anyone can choose the form of generality they care about most. But they must make clear: what is the object of the claim, what kind of change has been crossed, what effectiveness requirements apply, and which form of generality the existing evidence actually supports

.

A Second Question for Everyone #

In the previous essay, I wanted to leave everyone who cares about AGI with a first question: who is the subject of the AGI you are talking about? In this essay, I want to leave you with a second one: the General you are talking about is general across what kind of change?

This is also why we have spent so much effort taking General apart. The purpose is not to invent four new terms. It is that too many of today’s arguments about whether “AGI is almost here” are not actually discussing the same kind of General in the first place. Only after separating these different questions can we understand what a new capability breakthrough has really changed, and what questions it has still left unanswered.

In the future, when a laboratory, a researcher, or anyone else announces that “this system is already close to AGI” or “this model is already highly general,” we do not need to accept the claim immediately, nor do we need to reject it immediately. We can first place it against these four kinds of change: is it demonstrating skills across more tasks, transfer under unfamiliar conditions, task organization without a complete workflow, or attributable judgment by a persistent AI Actor when the rules cannot provide a unique answer? If the evidence demonstrates G1, let the conclusion stop at G1. If it supports G2, do not use it to imply G3. And if the system can plan autonomously, we still need to ask who actually formed the final judgment.

Only then can General stop being a term that anyone can interpret however they like and no one can truly verify. Only then can we begin to discuss seriously:

what kind of generality does the AGI we are talking about actually need?

In the previous essay, we first separated the question of AGI’s subject. In this piece, we have continued by separating the meanings hidden inside General. In fact, both essays are doing the same thing: before rushing to answer whether AGI has arrived, they ask us first to see clearly what question we are actually asking.

In my view, the genuine AGI we define must face an open, changing world that Humans can never fully script in advance. It needs broad capabilities. It needs to transfer under unfamiliar conditions. It needs to form its own tasks and paths. And sooner or later, it must also form judgment at points where existing rules cannot provide a unique answer. It is in this sense that G1 through G4 together open up the problem space through which we can understand General.

But even after we have clarified General in this way, another question remains. If a model becomes increasingly strong across G1, G2, G3, and even local instances of G4, does that mean AGI has already been established as a complete system? Capabilities can appear one by one, and at certain moments they can profoundly astonish humanity. But whether those capabilities already belong to the same persistent subject and whether they have been organized into a complete system that operates across time are still different questions.

That is where I want to continue in the next essay:

Why Is an Increasingly Capable Model Still Not Enough to Show That AGI Has Been Established?

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @kunyuan 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/02-what-does-general…] indexed:0 read:23min 2026-09-10 ·