Skip to content
Saved
Classification framework3 pillars, 9 inputs

The Robovations Score: one number, and every measurement behind it.

Nine inputs across three pillars, each scored 0 to 100 from the field record. Cost of ownership and data handling are published beside the score, not inside it.

Every score in the databaseOne column per point on the scale. Colour is the band. 382 robots carry a published score.
Dreame L40 Ultra · 90highest published scoreMedian 67Aiper Seagull Pro · 33lowest published score
0–3940–5455–6970–8485–100
0255075100

Published scores run from 33 (Aiper Seagull Pro) to 90 (Dreame L40 Ultra), with a median of 67. The hatched ends of the scale are unoccupied: nothing has scored above 90.

Deficient
3robots · 1%

Does not do what its category asks of it.

Limited
51robots · 13%

A person is routinely part of finishing the job.

Capable
188robots · 49%

Completes the job inside conditions the owner keeps within.

Strong
134robots · 35%

Does the job with known compromises, and the field record shows it holding up.

Exemplary
6robots · 2%

Finishes the whole job unattended and keeps doing it, with no pillar below the floor.

And the same scores by what the robot is forOne rubric runs across every category, so the numbers are comparable by construction. They do not land in the same place: robot lawn mowers sit at a median of 69, specialty robots at 57. Part of that is the machines and part of it is how much independent evidence each class has generated, which is why the gap is published rather than corrected for.
Robot Lawn Mowers88 scored69median
Robot Vacuums197 scored68median
Robot Pool Cleaners59 scored64median
Robot Window Cleaners19 scored61median
Specialty Robots18 scored57median
0255075100

Solid block is the middle half of the class, the tick is its median, the hairline is its full range. Categories with fewer than twelve scored robots are not plotted: a distribution over eight machines is three dots pretending to be a population.

Scale0 to 100Five bands, fixed across every category
Median robot67 · CapableHalf the database sits below this
Spread9.7 pointsStandard deviation across 382 scores
Anchored inField evidenceVendor material proves a commitment, never an outcome
The arithmetic

Where the hundred points go.

Three pillars carry the weight. Nine inputs carry the pillars. Each is scored 0 to 100 against its own written anchors.
40%

Capability

What it does without you.

Median 67Range 29 to 100
40%

Dependability

Whether it keeps doing it.

Median 67Range 13 to 99
20%

Ownership burden

What it asks of you.

Median 68Range 26 to 100
InputShare of the scoreHow it reads across the database
Capability40% of the scoreOwner footage and long-run reports, independent testing, and coverage of the failure modes documented for the category.

Task completion

Does it finish the job it is sold to do, start to end, with nobody stepping in.

16%40% of capability
median 7218 to 100

Coverage and navigation

Does it work the whole area systematically, and hold that coverage between runs.

10%25% of capability
median 7015 to 100

Edge-case recovery

When it meets the failure modes its category actually presents, does it recover and carry on.

10%25% of capability
median 5018 to 100

Operating envelope

How far conditions can vary before performance drops.

4%10% of capability
median 6825 to 100
Dependability40% of the scoreOwner reports at six, twelve and twenty-four months, warranty escalation threads, recall filings, parts catalogs and support terms with their dates.

Field failure rate

Documented hardware failures per owner-year, weighted by whether the failure stops the job.

16%40% of dependability
median 706 to 100

Software stability

Whether updates fix more than they break, and whether there is a channel where that can be seen.

14%35% of dependability
median 6515 to 100

Support longevity

What the maker has committed to for parts, service and updates, with dates attached.

10%25% of dependability
median 630 to 100
Ownership burden20% of the scoreManuals, owner upkeep diaries, consumable pricing histories, and third-party part availability.

Routine maintenance

Hands-on time the machine asks for in normal operation.

12%60% of ownership burden
median 6522 to 100

Consumables

Parts that have to be bought to keep it working, and how often.

8%40% of ownership burden
median 698 to 100

The bar on the right is that input across all 382 scored robots: the solid block is the middle half of the database, the tick is the median, the hairline is the full range.

No compensation

A strength cannot pay for a weakness.

Four rules run after the arithmetic and before the band. Each lowers a ceiling; none can raise a score. Where one fires, the robot page names it and prints the uncapped figure beside the capped one. The bar on each rule is solid where a score can still go, hatched where the rule closes it off.
  1. 1

    Every pillar at 70 or higher

    The top band asks for a machine with no weak side. One pillar below 70 stops the score at 84, however high the mean.

    84ceiling
  2. 2

    Every pillar at 55 or higher

    A Strong rating states that the field record holds up. A pillar under 55 is a documented soft spot, and the rating stops below the band.

    69ceiling
  3. 3

    Capability at 40 or higher

    Low ownership burden earns nothing when it follows from the robot doing very little. A machine with almost no capability has nothing to maintain and nothing to buy, and that must not read as an easy machine to own.

    54ceiling
  4. 4

    Support longevity at 60 or higher

    Below 60 the ceiling slides down by half a point for every point of shortfall, from 84 toward 54. A taper rather than a cliff, so a press release announcing a support policy cannot be worth more than a decade-long parts program.

    84 down to 54sliding ceiling

12 published scores are held down by these rules today. The rest govern a part of the scale nothing has reached.

Reported, not scored

Two things this number deliberately leaves out.

On these two axes more is not better, so both are stated on every robot page rather than graded into the number.

Five-year cost of ownership

$1,574median across 309 robots
$50middle half $1,050 to $2,250$7,040
  • Purchase priceWhat it costs to put in the house.
  • ConsumablesBrushes, filters, bags, blades, cartridges, on the cadence the manual sets.
  • SubscriptionAnything the maker charges for on a recurring basis to keep a shipped feature working.
  • Expected replacementBatteries and wear parts the field record shows failing inside five years.

A cheaper robot is not a better robot. The figure is published so a reader can weigh it against the score themselves.

Data posture

Local52
Core function needs no account and no network.
Cloud-optional42
Full core function without an account. The cloud adds convenience.
Cloud-linked273
Account required. Maps and telemetry leave the home, and core function survives an outage.
Cloud-dependent32
Core function stops without the vendor cloud.

A robot with no cloud is not a more capable robot. The posture is stated instead, next to what the manufacturer documents as leaving the home.

Coverage

When we publish no number.

A record we cannot measure is published as unmeasured, with the reason attached.
382ScoredNine inputs, each with an evidence line.
83Out of scopeNot sold through a consumer channel. Assessed under the Pre-Release Assessment instead.
19UnsourcedEnough of the rubric rests on sourcing we could not establish that a number would describe the record rather than the machine. The gap is named.
3IncompleteOur own assessment is unfinished. The inputs with no value are listed.
7Not yet scoredIn the database, awaiting an assessment pass.

Currently withheld

  • Degrii Zima ProOur assessment did not establish an independent source for 5 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate.
  • Piaggio gitaplusOur assessment did not establish an independent source for 7 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate, software stability, support longevity.
  • Degrii Zima LiteOur assessment did not establish an independent source for 5 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate.
  • Polaris AtlasSourced from Polaris product and parts pages only.
  • Betta SEOur assessment did not establish an independent source for 5 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate.
  • Tosima W5 Window Cleaning RobotOur assessment did not establish an independent source for 5 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate.
  • Seauto SAT40Our assessment did not establish an independent source for 5 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate.
  • Neakasa M1 PlusOur assessment did not establish an independent source for 2 of the nine inputs: software stability, support longevity.
  • Ecovacs Goat O1000 LiDAR PROOur assessment did not establish an independent source for 3 of the nine inputs: edge-case recovery, field failure rate, software stability.
  • Flymo EasiLife Go 200Under six months in this record and sourced only from Flymo.
  • Mamibot iGLASSBOT W168Our assessment did not establish an independent source for 7 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate, software stability, support longevity.
  • Polaris 9450 SportOur assessment did not establish an independent source for 4 of the nine inputs: coverage and navigation, edge-case recovery, operating envelope, software stability.
  • Pentair Prowler 930WOur assessment did not establish an independent source for 4 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope.
  • iRobot Roomba Max 715Our assessment did not establish an independent source for 9 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate, software stability, support longevity, routine maintenance, consumables.
  • Ecovacs Deebot T50S Pro OmniOur assessment did not establish an independent source for 6 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope, field failure rate, software stability.
  • Maytronics Dolphin Nautilus EON 100Our assessment did not establish an independent source for 4 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope.
  • Greenworks AiMowbot C20ZOur assessment did not establish an independent source for 2 of the nine inputs: coverage and navigation, edge-case recovery.
  • Greenworks Optimow 15Our assessment did not establish an independent source for 3 of the nine inputs: task completion, operating envelope, software stability.
  • LG Q9 AI AgentNo consumer retail channel for this machine has been established.
  • Maytronics Dolphin Oasis Z5iOur assessment did not establish an independent source for 4 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope.
  • Shark Detect ProOur assessment did not establish an independent source for 2 of the nine inputs: software stability, support longevity.
  • Husqvarna Automower 435 iQ AWDOur assessment did not establish an independent source for 4 of the nine inputs: task completion, coverage and navigation, edge-case recovery, operating envelope.
The parallel framework

Before it ships, there is nothing to measure.

The Robovations Score is built from field failure rates, update history, support commitments and consumable cadence, and none of those exist before a product reaches owners. A pre-release robot gets a Pre-Release Assessment instead, built from the signals that can be evaluated in advance. Lifecycle stage decides which of the two applies, and 83 records sit outside the Robovations Score today.
Lifecycle stageFrameworkWhat the page says
Pre-releaseAnnounced, not yet shipping to consumersPre-Release AssessmentPre-Release Assessment · 58 of 100
ProvisionalShipping under six months, thin owner dataRobovations ScoreRobovations Score · 61 of 100 · Provisional
VerifiedShipping six months or moreRobovations ScoreRobovations Score · 61 of 100
DiscontinuedWithdrawn from marketRobovations Score, frozenRobovations Score · 61 of 100 · Discontinued

The five dimensions

25%Track Record
Has this maker shipped what it demonstrated before, on schedule, with the capability it promised.
25%Engineering
How novel is the mechanical approach, and are the demonstrated capabilities achievable in shipping form.
20%Demo Match
What has been shown working, against what the marketing states. The size of the gap is the score.
15%Readiness
Announced date, price, distribution and regulatory status, or the absence of them.
15%Disclosure
How much of the product is actually disclosed. Fewer fundamental unknowns scores higher.

Its own band names

Deliberately different from the Robovations Score bands, so a reader knows at a glance which framework they are reading.

  • 0–29SpeculativeMinimal evidence and fundamental unknowns. The product may not be real.
  • 30–47Concept StageMeaningful gaps in plausibility or readiness. Direction is set, execution is not.
  • 48–63PlausibleA defensible path forward with material open questions.
  • 64–77PromisingStrong signals across most dimensions, with specific risks named and bounded.
  • 78–90Pre-Order ReadyHigh confidence the product ships as demonstrated. Narrow remaining uncertainty.
A pre-release assessment stops at 90.

The evidence that settles the question, whether the product ships and works in ordinary use, does not exist yet. No projection is published as a full mark, however strong the signals, and the remaining points are what shipping has to earn. Where the ceiling or a floor has moved a number, the robot’s own page names which one and shows the figure before it applied.

Three floors, so one strength cannot carry a record

Each names a dimension that decides the outcome on its own. Without them a company with no delivery record and no announced date could reach the top band on a persuasive demo reel, which is the claim this framework exists to resist.

  • under 60Track RecordCaps the assessment at 75, because the top band requires a maker that has shipped what it demoed before.
  • under 60ReadinessCaps the assessment at 75, because nothing is pre-order ready without a date, a price and regulatory evidence.
  • under 45Demo MatchCaps the assessment at 60, because the further what is claimed runs ahead of what has been shown working, the less the rest of the assessment is measuring the product that will actually ship.
The two never convert into each other.

When a pre-release robot ships, the assessment is archived, the first Robovations Score is computed from the shipping evidence, and the transition is logged on the tracker so a reader who followed the pre-release classification sees it change rather than discovering a new number. They measure different things, so the switch is exact rather than interpolated.

Prior art

What this score borrows, and where it departs.

Not derived from any single existing standard, though it takes one idea wholesale.
Euro NCAPVehicle safety ratings
Borrowed

The no-compensation principle. A car's star rating is limited by its weakest assessment area so a manufacturer cannot buy a rating in one place and neglect another. Our four floor rules are the same idea applied to a household machine.

Departs

Euro NCAP runs physical crash tests. We do not test. Our inputs are read from the field record, so our floors bound a measurement rather than certifying a result.

SAE J3016Driving automation levels
Borrowed

That autonomy is a spectrum with discrete behavioral breakpoints rather than a percentage. It anchors the Autonomy Ladder, which in turn anchors the capability pillar.

Departs

J3016 governs a vehicle on a road, where the operating envelope is regulated. A household is messier, and the safety floor is a different one.

ANSI/HFES 400Human readiness levels
Borrowed

That capability and readiness for real use are separate axes that should not be averaged together. Our Human Readiness Criteria run on a parallel track for the same reason.

Departs

HRL is a procurement framework for systems acquisition. Ours describes whether a household can adopt a robot today. The names are close, the intended reader is not.

ISO 13482Personal care robot safety
Borrowed

That safety, data handling and operational autonomy are distinct concerns that need to be assessed independently rather than rolled into one grade.

Departs

ISO 13482 is a conformity standard. The Robovations Score is a descriptive signal. We do not certify, and a robot scoring well here is not certified to anything.

Common questions

What readers ask about the number.

Is a 62 better than a 58?
On the measurements the score covers, yes, and the robot page shows which of the nine inputs the gap came from. It is not a purchase recommendation. Two robots at the same number can suit different households, and the score never accounts for what a specific home needs. Classification is not ranking.
Why nine inputs, and why are cost and privacy not among them?
Nine is the smallest set that covers the reads a household cannot verify before buying while keeping every input pointing the same way. Cost of ownership and data handling are reported beside the score instead. On those two axes more is not better: a cheaper robot is not a better robot, and a robot with no cloud is not a more capable one.
Why does nothing score 100?
A 100 would mean a machine with nothing left to demonstrate, and no consumer robot is there. What changed in revision 4.2 is what sits below the top: an anchor now describes the best product actually shipping rather than an imagined one, so a vendor that ships regular updates with no regression anyone has documented scores as the ordinary success it is instead of landing under halfway. The highest published score today is 90. The points above it stay unoccupied until a product earns them.
Do the band boundaries move to hit a target share?
No, and revision 4.2 is the test of that. The Exemplary band was sitting far under where it was expected to, and the boundary did not move — the anchors underneath it did, on the two inputs whose middle states described an absence of published reporting rather than a condition of the machine. Bands are set where the qualitative description changes. Moving a boundary to hit a number turns a criterion into a curve.
Can a robot score well and still be a bad buy?
Yes. The score measures the machine. Readiness measures the machine in the market, which includes availability, support and what the manufacturer has committed to. A well-scoring robot can still fail the readiness read if the maker has walked away from it or a recall is outstanding.
What counts as evidence?
Independent observation of behavior: owner reports and footage, teardowns, testing nobody involved paid for, regulatory filings. Manufacturer material proves a commitment, never an outcome, and a channel the vendor hosts, moderates, funds, or supplies the review units for counts as manufacturer material.
Can I see every score change over time?
Yes. Each robot page keeps a dated assessment history with the reason for each change. A framework revision is logged on the change log and triggers a rescoring of the whole database.
Next

See the nine inputs on a real robot, with the evidence line behind each one.

Dreame L40 Ultra
Last revised August 7, 2026. Applied to 382 scored robots.Suggest a correction