Live200 robots in operation across Europe as of May 2026.Live44 OEM partners and counting. Three new this month.Live11 European countries operational. Germany, Austria, Switzerland, France, Italy, Spain, Netherlands, Denmark, Sweden, Poland, United Kingdom.LiveFirst humanoid on Floor 2, Hamburg senior living. Week 12 of operation.PublishedCost-reduction case with a care group. Double-digit cost offset, year one.Live200 robots in operation across Europe as of May 2026.Live44 OEM partners and counting. Three new this month.Live11 European countries operational. Germany, Austria, Switzerland, France, Italy, Spain, Netherlands, Denmark, Sweden, Poland, United Kingdom.LiveFirst humanoid on Floor 2, Hamburg senior living. Week 12 of operation.PublishedCost-reduction case with a care group. Double-digit cost offset, year one.
werob.
Back to Magazine
Reading robot success rates honestly: what 45.7 percent floor pickup means for a shift plan
robot success rate

Reading robot success rates honestly: what 45.7 percent floor pickup means for a shift plan

Google published per-task success rates alongside its new models. They are strong research results and they are not shift ready for every task. Telling those apart is a numeracy problem, not a robotics problem.

werob· Systems integrator for robotics· 30 July 2026

Google published something unusual alongside the July 2026 robotics models: per-task success rates, task by task, hardware by hardware. Most vendors publish a video. The numbers deserve to be read carefully, because they say something more useful than either the enthusiasts or the sceptics take from them. The spread between tasks on identical hardware is far larger than the spread between good and bad robots, and that single observation should drive which use case you automate first.

Key Takeaways

The numbers, as published

These come from the Gemini Robotics 2 release, the vision-language-action model that converts perception into motor control. Note that this model is limited to early access partners, so the figures describe what the technology can do, not what you can buy this quarter.

Apptronik Apollo 2 with Inspire hands, whole-body manipulation

  • Pick up from a shelf: 76.3 percent
  • Pick up from a table: 68.4 percent
  • Pick up from the floor: 45.7 percent

Apollo 2 with SharpaWave hands, multi-finger dexterity

  • Unscrew a bulb: 92 percent
  • Tie a bin bag: 44 percent
  • Ziplock bag: 40 percent
  • Screw in a bulb: 36 percent
  • Dustpan: 32 percent

Franka Duo, gripper

  • Precise insertion: 89.6 percent
  • Tool kitting: 78.9 percent
  • General pick and place: 74.2 percent

The spread within one robot is the real finding

Look at the SharpaWave row again. Unscrewing a bulb works nine times in ten. A dustpan works three times in ten. Same robot, same hands, same model, same day. A factor of nearly three between two tasks that a person would describe as similarly easy.

The difference is geometry and contact. Unscrewing is a constrained rotation around a fixed axis with continuous feedback through the grip. A dustpan requires maintaining a precise blade angle against a floor while sweeping material that moves unpredictably, with the failure mode being invisible until the end. One task forgives small errors, the other accumulates them.

The same logic explains the height series. Floor pickup at 45.7 percent against 76.3 percent from a shelf is the same object and the same gripper. What changes is that a floor pick requires whole-body positioning, works against a cluttered background at an awkward angle, and offers less margin before contact. Choose tasks by their geometry, not by the robot's specification sheet.

The lever is the environment, not the machine

An independent data point sharpens this. The OK-Robot framework was evaluated in ten real homes across more than 170 objects and reported a total success rate of 58.5 percent on zero-shot pick-and-drop tasks. In cleaner, uncluttered environments the same system reached 82 percent.

Same software, same tasks, roughly 24 percentage points of difference, produced entirely by tidying up. For a facility operator this is the most actionable number in the entire field, because clutter is something you control and model weights are not. Defined drop zones, consistent container types, adequate lighting and clear floors move performance further than a hardware upgrade would.

This is also why demonstration videos generalise so poorly. They are filmed in the 82 percent world. Your loading bay at shift change is the 58 percent world.

What a per-attempt rate does across a shift

A rate per attempt is not a rate per shift, and conflating them produces business cases that fall apart in month two.

Take a task at 90 percent per attempt, run 200 times a shift. That is roughly 20 exceptions per shift. Each exception needs someone to notice it, walk to the machine, resolve it and restart the task. If that takes six minutes, you have consumed two hours of labour on a task that succeeded nine times out of ten. The success rate looks excellent. The staffing requirement is set entirely by the failures.

This inverts the usual evaluation. What matters is not the success rate but the absolute number of exceptions per shift and the cost of each one. A task at 74 percent that runs twenty times a shift is operationally calmer than a task at 95 percent that runs a thousand times.

It also explains why detection matters more than speed. An unnoticed failure in facility work costs far more than a slow success, because the follow-on effects, a missed delivery, an uncleaned area discovered by a customer, land somewhere else entirely. We treat that in our piece on progress detection as an operating metric.

Choosing the first use case

A workable rule of thumb from these figures, applied to the task rather than the robot.

Per-attempt rateReadingSensible next step
Above 85 percentCandidate for supervised operationPilot with a defined exception path
60 to 85 percentViable if the exception is cheap and quickly noticedPilot only where a person is nearby anyway
Below 60 percentNot a shift task yetChange the environment, or choose a different task

Notice that the honest recommendation for the bottom row is usually not a different robot. It is a different task, or the same task in a tidier environment. Transport between fixed points, repeated cleaning of open floor area and scheduled inspection rounds sit structurally in the top row. Picking varied objects off a cluttered floor sits in the bottom one, and no amount of procurement changes that this year.

werob specifies and integrates multi-vendor fleets for building operators, which regularly means recommending against automating a task. If you want to model the arithmetic for your own site, the calculator works from your numbers rather than ours.

FAQ

Is a 45.7 percent success rate bad?
For a research result on whole-body floor picking it is strong. As the basis for an unsupervised shift task it is not usable. The same model reached 76.3 percent picking the same object from a shelf, which tells you the constraint is the task geometry rather than the robot.
Which tasks are realistic to automate today?
Tasks with stable geometry and forgiving contact: transport between fixed points, repeated cleaning of open floor area, scheduled inspection rounds. The published figures put precise insertion at 89.6 percent and tool kitting at 78.9 percent on gripper hardware, while unstructured floor picking sits far lower.
How much does the environment matter?
More than most hardware decisions. The OK-Robot study reported 58.5 percent across ten real homes and 82 percent in cleaner, uncluttered environments using the same system. Defined drop zones, consistent containers, good lighting and clear floors are usually the cheapest available performance improvement.
Can I use these numbers in a business case?
Use them as a direction, not as a forecast. They were measured on specific hardware in a research setting, and the model they describe is limited to early access partners. For a business case, model the number of exceptions per shift and the cost of resolving each one, since that is what determines staffing.
Back to Magazine