A random sample measures your team's defect rate well and tells you very little about any one cleaner in any given week. One inspection of a cleaner who misses something on one job in ten catches them 10 percent of the time, and you would need 29 inspections of that person for a 95 percent chance of catching them. A realistic program gives you fewer than two per cleaner per week. So use sampling for what it is good at, measuring the team and aiming coaching, and if you need per-job certainty, the lever is not a higher sampling rate but making evidence a byproduct of every job.

The uncomfortable arithmetic of spot checks

Nearly every guide on this topic recommends inspecting a percentage of jobs. Almost none says what that percentage buys you, so here it is.

One note on whose side of the table this is written from. This page is for the cleaning company inspecting its own crews, where quality control is about defending an invoice, holding a contract, and knowing what to promise a client. If you are on the other side of that relationship, a property manager or host checking a cleaner you hired, the signals available to you are different and we cover them in how to check cleaners' work remotely.

If a cleaner leaves at least one thing wrong on a fraction of their jobs, and you inspect some number of their jobs at random, the chance you see the problem at least once is one minus the chance you miss it every time. That is the whole model. It is simple enough to be worth actually running, and the results are worse than most owners expect.

Chance a spot check catches a problem crew

Probability of seeing at least one defect, by the cleaner's true defect rate and the number of their jobs you inspect.

Cleaner misses1 job2 jobs5 jobs10 jobs20 jobs
1 job in 20 (5%)5%10%23%40%64%
1 job in 10 (10%)10%19%41%65%88%
1 job in 5 (20%)20%36%67%89%99%
About 1 in 3 (30%)30%51%83%97%>99%

Read the top-left corner slowly. A cleaner who leaves something wrong on one job in twenty, inspected once, looks perfect 95 percent of the time. For a 95 percent chance of having caught them at least once, you need 59 inspections of that one cleaner. At a 10 percent defect rate it is 29 inspections; at 20 percent it is 14.

Now put that against a real schedule. A company running 200 jobs a week with 12 cleaners, inspecting a respectable 10 percent of jobs, performs 20 inspections a week. Spread across the team that is fewer than two inspections per cleaner per week. Double the sampling rate to 20 percent, which is more than most companies can staff, and it is still only about three per cleaner per week. Against the table above, that is not a measurement system. It is a spotlight.

Read that as a limit on what a sample says about a person, not on what it says about your team. Those are different questions, and sampling is genuinely good at the second one. Twenty random inspections a week is roughly 1,040 a year, which pins a team-wide defect rate to within about two percentage points. Even a single month of that program gets you within about six. If the question is whether this quarter was better than last, or what rate to quote a client, random sampling answers it well.

What it cannot do is give you confidence about one specific cleaner on a short horizon, which is unfortunately the question most quality conversations are actually about. Your program can be simultaneously an accurate thermometer for the business and nearly blind to the individual, and most owners read the thermometer and assume it covers the individual too.

Two things also push real programs above the floor in that table, and it would be dishonest to leave them out. Inspections are rarely random: new hires, new properties, and anyone with a complaint on file get checked far more often, and targeted sampling substantially beats the numbers above. And defects are not independent across jobs, because the systematic misses are the persistent ones. A cleaner who never checks under the bed never checks under the bed, so the habit that matters most is also the easiest to catch. Treat the table as the floor for random sampling rather than as a description of a well-run program.

Two further caveats in the same spirit. Inspections are not the only detection channel you have; client complaints, review scores, and the next visit's condition all surface defects in parallel, and a sampling program sits on top of those rather than replacing them. And the model counts every defect equally, which no operator does. The established discipline here is acceptance sampling, where items are graded by consequence: a missing smoke detector or a safety issue warrants checking every job, while a dust ring on a baseboard genuinely can be sampled at a low rate. If you take one structural idea from this section, make it that one, because it is the cheapest way to spend a fixed inspection budget well.

Score against a published standard, not a homemade scale

The second thing most guides recommend is inventing a 1 to 5 scoring system. The problem is that two inspectors rating the same room with a homemade scale routinely disagree, which means the score measures the inspector as much as the work.

There is a published alternative. APPA, now styled Leadership in Educational Facilities and known until 2007 as the Association of Higher Education Facilities Officers, publishes a five-level cleanliness scale that institutions write directly into their custodial contracts. According to APPA's published cleanliness level standards, as reproduced in college facilities documentation (tctc.edu), the levels are described in observable terms rather than adjectives, which is what makes two people land on the same number. The University of Colorado Anschutz facilities group (cuanschutz.edu) states plainly that it targets Level 2 service, which is a useful reference point when a client asks what level they are buying.

APPA's five levels of clean

The published scale used across institutional facilities management.

1Orderly SpotlessnessShow-quality cleaning, developed for the corporate suite, the donated building, or the historical focal point.
2Ordinary TidinessThe level at which cleaning should be maintained. As Level 1, but up to two days of dirt, dust or stains.Recommended standard
3Casual InattentionThe first budget cut. A lowering of normal expectations, not yet unacceptable.
4Moderate DinginessThe second budget cut. Areas becoming unacceptable; the place always looks like it needs a spring cleaning.
5Unkempt NeglectThe lowest level. The facility is always dirty, with cleaning done at an unacceptable level.

Level 2 is the recommended standard, and Levels 3 and 4 are defined as budget cuts rather than as failures. APPA describes Level 3 as "the first budget cut" and Level 4 as "the second budget cut," which is the most useful idea in the whole scale: a drop in cleanliness is usually a staffing decision showing up on a surface, not a cleaner getting lazy. It also resolves an argument most cleaning companies have internally, which is whether the target is perfection. It is not. Level 2 tolerates two days of ordinary dust, and pricing a Level 1 outcome into a Level 2 contract is how crews end up running long on every job.

The honest caveat: APPA was written for institutional buildings measured in cleanable square feet, not for a three bedroom house or a turnover between guests. Do not adopt the audit forms wholesale. Adopt the vocabulary and the target, because having a shared, externally defined name for "good enough" is worth more than the precision of any scoring sheet, and because a client argument goes differently when the standard is somebody else's published scale rather than your opinion.

What visual inspection can and cannot establish

Everything above assumes an inspector looking at a room, which is the method essentially every residential and turnover cleaning company uses. It is worth knowing what gets said about that method, and by whom. Writing for ISSA (issa.com), the cleaning industry association, a technical writer at ATP-instrument maker Hygiena describes visual inspections as "imprecise, subjective," and potentially unacceptable in many facilities, especially healthcare. That is a vendor making a case for instruments, so weigh it accordingly, but the underlying point survives the conflict of interest.

The objective alternative used in institutional and healthcare settings is ATP testing, which swabs a surface and reports organic residue in relative light units. ISSA is careful about what it proves: ATP monitoring does not identify bacteria or viruses directly, it detects the general presence of organic matter that microbes can use to grow. Higher readings mean greater potential contamination rather than a specific pathogen count.

For a residential or short-term-rental cleaning company, ATP meters are almost always the wrong tool. Your client is not a hospital and your standard is not microbiological. The useful takeaway is narrower and more actionable: even the case made by instrument vendors concedes that looking at a room is a subjective measurement. If your entire quality system rests on one supervisor's eye applied to a small sample of jobs, you have a subjective measurement applied to a sample that the arithmetic above says is too small to be conclusive. Both weaknesses point the same direction, and it is not toward buying instruments.

A quality program that actually holds

The fix is not a higher sampling rate, because the table shows how quickly that gets expensive for how little it buys. The fix is to make evidence a byproduct of doing the job, so that reviewing work stops requiring a trip and coverage stops being a budget question.

  1. Write the standard down, and version it

    Name your target level and define it per room in observable terms. Version the document with a date, because the most common dispute in cleaning quality control is two people applying different editions of the same checklist. When a standard changes, the change should have a date attached so you can tell what applied on the day of the job.

  2. Make completion photos a fixed set, not a free-for-all

    Require the same views on every job of a given type: the same kitchen angle, the same bathroom angle, the same floor shot. Fixed views are what make comparison possible, both against the standard and against the same property last time. A pile of unrepeatable photos proves someone had a phone out.

  3. Capture in-app so the evidence carries a timestamp

    Photos taken inside your ops tool arrive with a capture time attached. Photos uploaded from a camera roll do not, and reuse of a previous visit's photos is the most common way completion evidence goes wrong. We covered the detection side of that in our piece on whether cleaners can fake cleaning photos.

  4. Track failures by item, not only by person

    Count how often each checklist item fails across the whole team. An item failing across several cleaners is a standard problem or a time problem, not a people problem, and disciplining individuals for it burns trust and fixes nothing. Item-level data is also what tells you which parts of the job are genuinely hard.

  5. Reserve in-person inspection for what photos cannot show

    Smell, humidity, noise, the feel of a floor, and anything requiring a hand on a surface still need a person. That is the right use of your limited inspection budget, rather than spending it re-checking things a photo already settles.

  6. Close the loop with the cleaner, with the evidence in front of you

    A miss discussed the same day against a photo and a written standard is training. The same miss raised a week later from memory is an argument. Speed matters more than severity here, and the entire value of the evidence trail is that it makes the conversation specific.

Steps two and three are the ones that change the math. Once every job produces a comparable, timestamped set of views, you are no longer sampling jobs. You are sampling the same fixed views inside every job, which trades an unknown miss rate across jobs for a known blind spot inside each one. That is a better trade, not a solved problem: step five exists precisely because a photo set cannot show you smell or a sticky floor. What changes is that the question stops being how many jobs you can afford to inspect and starts being how quickly someone can look at what every job already produced.

Quick FAQ

What percentage of cleaning jobs should you inspect?

It depends what you want the sample to do. To measure your team's defect rate, a modest random sample is plenty: 20 inspections a week is about 1,040 a year, which pins the rate to within roughly two percentage points. To be confident about one specific cleaner, no realistic percentage is enough. Inspecting 10 percent of jobs across a 12-cleaner team running 200 jobs a week is under two inspections per cleaner per week, while catching a cleaner who misses 1 job in 10 with 95 percent probability would take 29 inspections of that one person. Sample randomly to measure, target extra inspections at new hires and complaint history to detect, and check safety-critical items on every job rather than sampling them.

What standard should a cleaning company score against?

Use APPA's five levels of clean rather than inventing a 1 to 5 scale. The levels run from Level 1 Orderly Spotlessness through Level 2 Ordinary Tidiness, Level 3 Casual Inattention, Level 4 Moderate Dinginess, and Level 5 Unkempt Neglect, and APPA recommends Level 2 as the reasonable standard for normal operations. It was written for institutional facilities rather than homes, so it works as shared vocabulary and a target rather than as a residential scoring sheet, but it beats a homemade scale because the level descriptions are observable and two different inspectors tend to land on the same number.

Is visual inspection good enough for cleaning quality control?

It is the standard method and it has known limits. ISSA describes visual inspections as imprecise and subjective, and notes they may be unacceptable in some facilities, particularly healthcare. The objective alternative used in institutional settings is ATP testing, which measures organic residue in relative light units, though ISSA is clear that it detects organic matter generally rather than identifying bacteria or viruses directly. For a residential or turnover cleaning company, ATP meters are usually overkill; the practical upgrade is not a better instrument but broader coverage.

How do you inspect cleaners' work without being on site?

Make the evidence a byproduct of the job rather than a separate errand. Require completion photos of the same fixed set of views on every job, captured in-app so they carry a timestamp, and pair them with a versioned checklist so you know which standard applied on that date. That converts inspection from a trip someone has to make into a review someone can do from anywhere, which is what makes full coverage possible instead of sampling.

How should you handle a cleaner who fails an inspection?

Separate the miss from the pattern. A single failed item on one job is a coaching conversation held against the written standard, ideally with the photo in front of both of you. A repeated miss of the same item is a training or process problem, not a discipline problem, and usually means the checklist is ambiguous or the time allowed is short. Track failures by item rather than only by person, because the same item failing across several cleaners points at your standard rather than at them.

Sources

Sources are named at the publisher level with their root domain, rather than linked or titled; every figure is verifiable at the named source.

  1. APPA cleanliness level standards, published in college facilities documentation, Tri-County Technical Collegetctc.edu
  2. Custodial cleaning levels and APPA service target, University of Colorado Anschutz Facilities Managementcuanschutz.edu
  3. Vendor-authored guidance on inspection methods and ATP monitoring, published by ISSAissa.com
  4. Sampling probability calculations, RapidEye Researchrapideyeinspections.com

Keep reading