Physical AI aims to connect intelligence with real-world action.
A user might say:
"Bring me the bottle from the kitchen."
A robot must turn that high-level instruction into a sequence of grounded actions.
Natural Language
|
v
Task Understanding
|
v
World Model
|
v
Task Planning
|
v
Motion Planning
|
v
Control
|
v
Physical Robot
The important insight is that language understanding alone is not enough.
Consider:
"Pick up the bottle."
The system must identify:
Therefore:
Language
+
Vision
+
Robot State
+
Environment Model
|
v
Grounded Action
A foundation model can produce structured actions rather than motor commands:
{
"action": "pick",
"object": "bottle",
"location": "kitchen_counter"
}
The robotics stack then translates this into navigation and manipulation primitives.
A high-level instruction can be decomposed:
Bring bottle
|
+--> Navigate to kitchen
|
+--> Find bottle
|
+--> Reach bottle
|
+--> Grasp bottle
|
+--> Navigate to user
|
+--> Release bottle
Each subtask can be executed and verified independently.
/natural_language_task
|
v
/task_planner
|
v
/world_model
|
v
/action_executor
/ v v
/navigation /manipulation
Physical AI should use closed-loop execution:
Plan
|
v
Execute
|
v
Observe
|
v
Verify
|
+---- success ---> Next Step
|
+---- failure ---> Replan
This is critical because the physical world is uncertain.
A grasp may fail. An obstacle may move. A door may be closed.
Foundation models should operate behind explicit constraints:
Separate responsibilities:
Foundation Model
|
| high-level intent
v
Task Planner
|
| structured actions
v
Robot Skills
|
| validated commands
v
Motion Planner
|
v
Controller
This makes the system easier to test and replace.
Evaluate both intelligence and physical execution:
The future of physical AI is not simply putting a large model inside a robot. It is building a reliable bridge between language, perception, world models, planning, and safe physical control.