bytevyte
bytevyte
Language
ai-beats —

Figure Ties Humanoid Robot Training Data to a Sixfold Dexterity Gain

humanoid robot training data

Figure has reported that its Helix 2.5 model completed household chores at a 56% success rate inside 30 Bay Area homes it had never been trained in. An otherwise equivalent model trained without the company's Index pretraining managed 9%, which makes the gain a more than sixfold improvement. Figure credits a single changed variable: the volume of humanoid robot training data fed into the model before it entered a home.

The result arrived in late September 2026, roughly four weeks after Figure committed billions of dollars to acquiring physical-AI data. It is the most concrete public evidence yet for the company's thesis that data scale, not model architecture, is what limits humanoid dexterity. That thesis now underwrites a large share of the capital flowing into industrial robotics, and it still rests on numbers Figure produced and Figure reported.

The three tasks were long-horizon domestic chores: tidying a living room, making a bed, and folding towels. Figure measured full-task completion rather than partial credit, across homes and objects the robot had not encountered during training. The company describes the 56% figure as zero-shot success, meaning the model transferred to unfamiliar environments without task-specific fine-tuning.

The Experiment Behind the Number

Figure presents the comparison as controlled. The same Figure 03 hardware, the same task data, and the same Helix vision-language-action architecture were held constant across both conditions. Only the pretraining input changed. One model received the Index corpus; the other did not. Figure's argument is that this isolation is what makes the sixfold gap attributable to data volume rather than to a hardware revision or a new model generation.

Figure reported a second result that may matter more than the headline. Increasing Index pretraining data across four runs produced predictable improvements in downstream robot-action prediction. A monotonic relationship between more data and better action prediction is the signature of a scaling law, and it is the kind of curve that turns a benchmark into an investment case. If the improvement is repeatable and forecastable, money spent on data carries a defensible expected return.

The scaling behavior is not unprecedented. Prior work on human action pretraining, including Dyna's, had already shown that error rates fall according to a power law as pretraining volume grows, extending from pretraining scale through to closed-loop real-robot performance. Human action pretraining has been demonstrated at 1.2 trillion tokens. Figure's contribution is to extend that predictable scaling to full-body humanoid systems and to attach it to real households rather than to a laboratory setting.

Figure has been careful to bound its own claims. The company states that general humanoid robotics is not solved, and it does not present the chore result as proof of factory readiness.

Index and the Cost of Humanoid Robot Training Data

The data behind the result comes from Index, a gig marketplace Figure launched earlier in 2026. Human workers, paid as creators, film themselves performing domestic tasks, and Figure converts that footage into training material for its robots. Figure has said it paid $15 million to creators globally during the platform's launch phase.

The mechanics set the cost structure. Building a humanoid robot training data pipeline means paying a distributed workforce to generate footage, curating it, and converting it into labeled action sequences at scale. That is an operating expense that compounds with volume, and it explains why Figure framed its commitment in billions rather than millions. The 56% result is the first measurable return on that spend.

Figure is not alone in reading the problem this way. At IEEE Humanoids 2026, factory labour dominated the agenda, and the data-centric framing recurred across competing programs. Google DeepMind's Gemini Robotics models are being paired with Boston Dynamics Atlas hardware, while Apptronik's Robot Park is designed to feed real production task data back into model training. Apptronik, Boston Dynamics, and Tesla are all shifting from pilot deployments toward industrial adoption, where buyers ask for measured task success instead of staged demonstrations.

That convergence cuts both ways for Figure. A shared thesis validates the strategy and de-risks the category for investors. It also means a data pipeline is not a durable moat on its own, because any well-funded competitor can rent the same gig workers and film the same chores. The durable advantage, if it exists, comes from volume and iteration speed rather than from a proprietary method. Figure's $39 billion valuation, set in late 2025, already prices in an assumption that this lead is real and defensible.

Buyers in manufacturing and logistics have been explicit that they will pay for measured throughput rather than capability demonstrations. Humanoid developers have spent years showing that bipedal machines can walk, balance, and manipulate objects; the open problem is reliability in unstructured settings. A metric that improves because of more data is easier to forecast than one that improves because of a cleverer model, which is why the data-centric narrative has drawn capital so readily.

The capital cycle depends on this claim in a specific way. Humanoid developers have raised money against a future in which general-purpose robots perform useful work at scale, and the value assigned to that future falls sharply if dexterity improves only slowly. Data-centric scaling offers a path where more spending produces more capability, an easier story to finance than a research program whose next breakthrough may or may not arrive. Figure's four-run curve is the kind of evidence that keeps that financing open.

What the 56% Does Not Prove

The headline figure is self-reported, and it measures chore success rather than validated factory throughput. Those are different quantities. A robot that completes 56% of household tasks in a domestic setting may behave very differently on a warehouse floor with reflective surfaces, partial occlusions, and unfamiliar packaging, where success rates are known to collapse when the training distribution does not match the deployment environment.

A 56% full-task success rate also means the robot fails more often than it succeeds. That is far below the reliability threshold for unsupervised operation, and it is the gap that separates a research demonstration from a product a customer will pay for.

Independent replication is the missing piece. No third party has reproduced the sixfold gain under its own protocol, and the comparison is against a deliberately weakened baseline: a robot trained from nothing rather than against a strong competing system. The fair reading is that Figure has demonstrated its pretraining method outperforms a baseline with no pretraining. It has not shown that it outperforms its rivals.

Bimanual manipulation raises the bar further. Coordinated two-handed tasks require training data that captures hand-to-hand coordination, contact dynamics, and recovery behavior, not just end positions. Whether the Index corpus captures those properties well enough to generalize beyond the 30 homes is unresolved.

What would change the picture is a competitor matching the result by a different route. If a rival reaches comparable household success through better architectures or simulation-heavy training rather than through rented human footage, the data pipeline stops looking like the decisive input and starts looking like one option among several. No such head-to-head comparison exists in public, which is why both the believers and the skeptics can read the same 56% figure and reach opposite conclusions about where the industry is heading.

Why this matters

The result matters less as a product milestone than as a test of where value in humanoids sits. If data volume is the binding constraint, the companies able to fund and operate large humanoid robot training data pipelines hold the advantage, and hardware differentiation counts for less than the balance sheet behind the data. Figure has now put a number on that claim, and its peers building their own data flywheels are betting the same way.

The number to watch is not 56%. It is whether Figure publishes a larger run that holds the curve, and whether an outside lab reproduces the gap on its own hardware. Until then, the sixfold jump is a strong internal result rather than a settled law of robotics.

✔Human Verified


Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.