The cost of teaching a robot something new
Before paying for another thousand hours of robot demonstrations, what evidence should a team demand that those hours will help?
By Clairevue · · 6 min read

Cobot is recruiting more than 30 people in Santa Clara to work with robotic arms and collect training data. The shifts run from 7 a.m. to 3 p.m. and from 3 p.m. to 11 p.m. A company that wants those hours filled is making a substantial commitment to teaching machines through human demonstrations.
Its founder, Brad Porter, had a reservation. In a recruiting post we found through our Physical AI Signals feed, he wrote that scaling collection without a clear training recipe and evidence of scaling laws was “a good way to burn a lot of money not getting to the destination.” Then he added: “We’re scaling data collection now.” [1]
Porter doesn’t disclose the results behind that decision. But his warning raises a question that every team buying robot demonstrations has to answer: what evidence would justify paying for another thousand hours?
A useful answer starts with a capability the robot needs to improve and a test of whether the new experience helps it somewhere unfamiliar. Otherwise, a busy collection room can produce an impressive archive while the robot keeps making the same mistakes outside it.
What a collection shift buys
Porter has experience with the work of deployment. Cobot’s Proxie moves wheeled carts through workplaces; at its November 2024 unveiling, The Robot Report described trials with customers in logistics, healthcare and other settings. [2] That gives his concern about collecting without a learning recipe some practical weight. A customer needs a robot that can do a job, and someone has to work out which training investments bring it closer.
The new hiring effort focuses on robotic arms. Those operators will work through tasks and scenarios, producing records a model can learn from. Making those records useful takes coordination as well as time at the controls.
At Deccan AI, a supplier of AI training data, that coordination appears in an operations vacancy in Hyderabad. Delivery leader Syed Shakeel Imdad shared the role in a post about the fieldwork behind physical-AI data collection. [3] The job description calls for managing field teams, vendors, annotators and quality-control staff, allocating workers, investigating problems and visiting collection sites. [4]
The vacancy tells us what a collection program has to organize. It says nothing about how much any customer's robot improves. To decide whether the effort is worthwhile, an operations team needs a more specific brief than “collect more.”
Imagine a robot that can put a cup into a drawer in the collection room but struggles with drawers elsewhere. An operator could repeat the familiar task for another week. The team might instead collect demonstrations with different handles, drawer heights and objects in the way. Both plans consume shifts. They pursue different explanations of the failure.
That example is hypothetical, but the choice is concrete: should the next batch deepen practice on a familiar setup, or broaden the situations the model encounters? Research from Physical Intelligence offers a way to investigate it.
What a different kitchen teaches
Physical Intelligence’s April 2025 paper on its π0.5 model studied household tasks in environments the robot hadn't seen during training. The researchers wanted to understand how a learned skill could carry over to another home. [5]
Their recipe included about 400 hours of mobile-robot demonstrations from roughly 100 home environments. It also drew on other robots, laboratory tasks, web data and annotations that divided activities into smaller steps. The model could learn to select a step such as “pick up pillow” and then produce the movements to carry it out.
One experiment varied the number of locations represented in the mobile-robot training data: 3, 12, 22, 53, 82 and 104. The researchers used a shared pre-training setup, then compared post-training runs with the same training-step budget and the same number of unique samples seen. That helped separate the effect of seeing more environments from simply consuming more examples.
On their held-out mock-home tasks, performance generally improved as the number of training locations rose. Those tasks included putting dishes in a sink, packing items into a drawer, moving laundry into a basket and making a bed.
For a collection team, the implication is useful and specific. Under this recipe, changing where demonstrations came from improved performance somewhere new. Location diversity was a property they could deliberately buy with their collection budget and evaluate afterward.
The researchers also tested which parts of the data mixture mattered. Removing demonstrations from other robot platforms reduced performance. Removing web data had a more selective effect: it hurt performance with unfamiliar object categories and high-level task decisions, while its effect on the full mock-home tasks was not statistically significant in the reported comparison.
So the next collection request depends on the weakness being addressed. A model that doesn't recognize an unfamiliar object has a different problem from one that recognizes it but cannot manipulate it. These experiments suggest that the sources of experience can contribute differently to those capabilities.
An illustrative workflow for planning and testing a collection strategy. Diagram: Clairevue; not a company's system architecture.
This is evidence for a particular training approach, rather than a universal rule for buying robot data. The final real-home evaluation covered three previously unseen homes. Some task scores measured the fraction of steps completed: moving half the dishes could earn partial credit, which is different from finishing the whole job. Those limits matter when deciding how much confidence to place in the result.
Even with that boundary, the paper supplies something the hiring posts cannot: experiments connecting choices about training data to measured changes in performance.
Before booking the next thousand hours
A collection plan should make that connection explicit before the shifts are booked. What does the robot currently fail to do? What will the new demonstrations change? Where will the team test whether the skill improved?
For the hypothetical drawer problem, a team could collect across more rooms, train a candidate model and compare it with the previous version on drawers excluded from collection. It should count completed tasks and any human help, so partial progress doesn't get mistaken for reliable completion. If the new model still struggles, that result should inform the next collection request.
The spending decision then rests on a comparison: does a model trained on the new batch complete more of the same unfamiliar tasks, with less human help, than the previous version? Holding the test conditions steady gives the team a way to judge what the extra collection bought. It also gives them a reason to change course when a batch adds little.
The operators arriving at 7 a.m. can generate more demonstrations. Their time becomes a defensible investment when the team can explain why those demonstrations are the next ones the robot needs—and show afterward that they helped.
Sources
- Brad Porter / Cobot: Recruiting post on data collection and scaling.
- Mike Oitzman, The Robot Report, November 20, 2024: Collaborative Robotics unveils Proxie mobile manipulator.
- Syed Shakeel Imdad / Deccan AI: Field-operations hiring post.
- Deccan AI: Associate Program Manager – Physical AI Data Operations.
- Physical Intelligence, April 2025: π0.5: a Vision-Language-Action Model with Open-World Generalization, v1, particularly sections IV-C, V-A, V-B and V-C. A documented research case, not a claim about the latest state of the field.