Watch a robot pick a ripe tomato without crushing it, and you are really watching a recording play back through a machine. Somewhere, a person performed that motion first, carefully enough that a model could learn its shape. This is how most robotic manipulation gets built today. Not from rules typed into software, but from demonstrations captured off real people and turned into training data. Following one demonstration from start to finish shows why the process is so labor-heavy, and why it is not going away.
It starts with capture. The raw material here is what the field calls human demonstration data for robotics: the hand approaching an object, the grip forming, the small correction when the weight comes out wrong. Demand for it has climbed with the sector’s funding. Robotics startups pulled in $18.8 billion globally in 2026 so far, more than in all of 2025, according to Crunchbase, and every new robot design needs its own examples. A model cannot generalize from a task it has never once watched a human perform.
Why the internet has nothing to offer here
The obvious question is why this data must be made at all. Language models, after all, learned from text that already existed in vast supply. The trouble is that physical skill leaves no written record. Nowhere on the web is there a usable log of how much pressure it takes to hold an egg, or how a wrist rotates to seat a plug in a socket. That knowledge lives in bodies, and it has to be recorded on purpose, sensor by sensor. It is the main reason a field awash in capital still fumbles chores a small child manages without thinking twice.
Capturing the motion is only the first half of the job. Before a model can use the footage, it has to be labeled, and one form of labeling has become central to robotics. egocentric video annotation works on first-person recordings shot from the doer’s point of view, close to what a robot’s own head camera sees, and marks the details that matter: where the gaze lands, which object the hand selects, the exact frame where contact begins. Meta’s Ego4D dataset gathered 3,670 hours of this first-person video, and its follow-up, Ego-Exo4D, took more than 200,000 hours of annotator effort to label.
The scale hiding inside “just label it”
Those Ego4D numbers are worth sitting with, because they show how the labeling load dwarfs the capture. Grand View Research valued the broader data collection and labeling market at $3.8 billion in 2024 and projects it to reach $17.1 billion by 2030, growing at nearly 28 percent a year, with image and video the largest slice. That growth is not marketing noise. It reflects a stubborn ratio: a single hour of raw demonstration can take several hours to annotate properly, and a model often needs thousands of well-labeled hours before it becomes dependable. Multiply that across every object a home robot might touch, and the labeling bill grows faster than the collection bill ever could.
What good data actually looks like
Not every demonstration earns its place in a dataset. The examples worth keeping include the ones that go wrong, because a robot that has never seen a recovery cannot perform one when its own grip slips. Coverage matters just as much: different lighting, different objects, different starting positions, gathered until the model stops being surprised by the ordinary. The International Federation of Robotics counted 542,000 industrial robots installed in 2024, lifting the global fleet in operation to about 4.66 million. Each new machine, and each new task it takes on, tends to demand its own fresh round of demonstrations rather than borrowing another robot’s. Transfer between different robots helps at the margins, but it rarely removes the need for fresh, machine-specific capture.
The people behind the autonomy
By the time a robot performs a task smoothly on its own, dozens of people have already touched the data behind it. The operators and volunteers who performed the motions. The annotators who labeled every relevant frame. The reviewers who threw out the runs that were too sloppy to trust. That human layer does not vanish as the models improve, which is the part outsiders tend to miss. Each new gripper, task, and environment sends the entire cycle back to the beginning.
This is the uncomfortable shape of progress in physical AI right now. The promise is machines that spare people repetitive work, yet building them depends on a growing amount of exactly that kind of work, done by hand and rarely credited. The robot in the clip looks like it figured the task out for itself. In truth, people taught it, one recorded demonstration at a time, and for the foreseeable future that is the only way the teaching gets done.
