Physical AI & Robotics · Open-access guide

Robot Task-Success Detection: Validating Automated Labels

Validate automated robot success labels with physical acceptance evidence, task-level errors, blinded audits and checks after policy or sensor changes.

Stroncature Research · Sources checked · Editorial method

Validate automated robot task-success labels against the required physical outcome, using previously unseen task episodes and independent evidence. Measure false acceptance and false rejection at the expected proportions of successful and failed tasks, then audit training labels. A benchmark average cannot establish that a visual judge detects contact, quality or safety conditions reliably in another workflow.

Physical acceptance rules and detector errors

A robot’s apparent completion can differ from the state a customer requires. A connector may look aligned without being seated; an object may reach the destination with damage that a distant camera does not reveal. Before evaluating a model that labels outcomes, define the task’s acceptance rule and the observation needed to verify it. Some rules can be supported by images, while others require force, dimensional, electrical or other process evidence. The evaluator cannot recover information that its input does not contain.

The FailBench preprint, released in September 2026, evaluates detectors across multiple public manipulation datasets and reports marked difficulty on contact-intensive tasks. Its relevance is the variation across sources, not a universal industrial error rate. Curated benchmark prevalence, camera views and task labels differ from production. A result can identify a weakness worth testing while remaining unsuitable as a direct estimate of the number of defective parts that would escape inspection at a particular site.

Distinguish the consequences of the two principal errors. A false success can admit a failed trajectory into positive training data or permit defective output to continue downstream. A false failure can discard useful experience or cause an unnecessary retry. Overall accuracy combines those outcomes, although their costs may be very different. Preserve task-level error counts, the number of true successes and failures, and the uncertainty associated with a limited sample. A model that performs well on obvious object movement may still be unsuitable for a subtle assembly condition.

Failure prevalence and independent label audits

Base rates change the meaning of an alarm. In an illustrative set of 10,000 episodes with 100 genuine failures, a detector that catches 90 failures and incorrectly flags 5% of the 9,900 successes produces 585 failure flags. Only 90, approximately 15.4%, are genuine failures. The arithmetic does not describe any benchmarked model. It shows why a useful failure-detection rate can still generate an expensive review queue when successful episodes dominate ordinary operation.

Build an audit set independently of the automated verdict. Human reviewers should not see the judge’s answer before making their own assessment, and ambiguous contact outcomes should use the relevant process measurement rather than visual majority opinion. Hold out meaningful changes in camera, robot, site and task when testing transfer. RH20T’s multimodal collection illustrates that robot datasets can contain information beyond ordinary video. Whether an additional signal helps a given acceptance rule still needs controlled evaluation.

Training use, model changes and verification cost

The intended use determines the required validation. A judge that triages obvious clips for research can be useful even if ambiguous cases need people. A reward model shapes the behaviour a policy learns, making exploitable errors more consequential. A final quality decision may require independent deterministic checks. Keep those uses separate instead of adopting a single threshold everywhere. Record which episodes were automatically accepted, escalated, overruled or confirmed, and preserve enough provenance to correct the training set when errors are discovered.

Validation must continue after the judge enters a learning loop. Retraining changes the policy’s behaviour and therefore the distribution the judge sees. A policy can learn to satisfy a visual proxy without meeting the physical task. Repeat blinded audits after material policy, camera, tooling or judge changes, retaining a frozen comparison set and newly collected representative episodes. Version the input, model and acceptance rule together so a later improvement does not overwrite the evidence of how previous labels were produced.

The economic comparison includes inference, data handling, instrumentation, review and the consequence of residual mistakes. Cheap labels become valuable when they reduce the complete cost of obtaining trustworthy training or evaluation evidence. A large automatically labelled corpus is not inherently a stronger asset than a smaller outcome-confirmed one. The useful measure is verified information that improves accepted robot performance, with uncertainty and correction costs preserved throughout the process rather than hidden behind an aggregate accuracy figure.

Email newsletter

Physical AI Finance Monitor

Physical AI Finance Monitor follows robot-learning evidence and data economics, examining whether larger datasets and automated evaluation translate into reliable deployed performance.

Sign up for the free newsletter

Newsletter sign-up is free. Access to paid reports depends on the subscription selected.

About this publication