Factory AI

Beyond Sampling: How to Start Computer-Vision QC Without It Failing

Vision-based inspection projects fail for a short list of recurring reasons, almost never because the model was not good enough. Here is what actually goes wrong and a safe order of operations.

"We inspect a 5% sample."

It sounds like an orderly process. Translated literally, it means the defects that reach your customer are a matter of probability, not something you control.

Sampling is not the wrong method. It is the most sensible method available under one constraint — inspection requires human eyes, and people are finite in number, get tired, and are less accurate at three in the afternoon on the second shift.

That constraint has changed. Yet a great many projects trying to take advantage of the change fail, and almost never because the technology was not good enough.

Why most projects fail

1. Starting from the technology instead of from the most expensive defect

The common pattern: an executive sees a demo and instructs the team to "use AI for QC". The team picks the easiest place on the line to install, which is usually not where the problem is.

Six months later the system works, the accuracy numbers look good, and nobody feels anything has improved — because the place that actually costs money was never touched.

What to do instead: start from the question "which type of defect cost us the most money last year?" The answer is usually not the most frequent type. It is the type that reached a customer and came back as a claim, because the cost of a defect caught on the line and one that escapes differ by a large multiple — return shipping, replacement, the sales team's time, and damage to a relationship that cannot be priced.

2. No images of defects for the system to learn from

This is the real obstacle in most plants, and it sounds contradictory.

A plant with good quality control has very few defects, which means very few examples for a system to learn from. Some types occur once or twice a month, and nobody photographs them, because when one appears you simply pull it out and carry on.

So when the day comes to start a project, "could we have some images of the defects" is answered with "we do not have any".

What to do, and to start today even if no project is planned: put an ordinary camera on the reject station and photograph every piece pulled out, recording which defect type it is. No system is needed. A phone and folders named by defect type is enough.

Accumulate six months of that and you have the most valuable thing this kind of project can have. Money cannot buy it, and it determines success more than the choice of vendor does.

There is another route for cases where defects genuinely are rare: teach the system what good looks like in detail, and have it flag anything that departs from it. That needs a large number of images of good pieces, which are much easier to come by, at the cost of over-alerting early on.

3. Setting an accuracy threshold without comparing the real process

If a team sets a threshold without defining defect types, counting rules, risk, and human roles, the trial may not produce evidence that supports a decision.

A more careful frame: compare the proposed system with the current inspection process using data from the same line. Define coverage, false accepts, false rejects, error cost, and human-review points. Illustrative figures from another line should not be treated as a guarantee or substituted for a real trial.

4. The system decides, gets it wrong, and nobody trusts it again

The fastest way to kill a project is for the system to reject hundreds of good pieces because the lighting changed, or to pass a whole batch of defective ones. After that nobody in the plant believes it again, and rebuilding that trust takes longer than starting over.

What prevents it: the system needs a confidence threshold. Pieces it is not confident about go to a person, rather than being guessed at and acted on.

The principle is the same one that applies to document reading — a system that knows what it does not know is worth more than a more accurate system that is confidently wrong every time.

5. Nobody in the plant owns it

The vendor installs the system. It works well for three months. Then the part specification changes, or a light gets moved, and it starts flagging incorrectly. Nobody on site can adjust it, so they wait for the vendor, switch it off in the meantime, and never switch it back on.

What to put in the contract: at least one person in the plant must be able to adjust thresholds, review results, and add new examples for the system to learn from. If every change has to go through the vendor, the system dies within the first year.

A safe order of operations

Step 0 — Collect images (start today, costs nothing)

Photograph every rejected piece, in folders by defect type, recording the date and the shift. Keep it up for at least three months before talking to any vendor.

This step puts you in a completely different negotiating position, because you will understand the real problem before the vendor does.

Step 1 — Pick one point

Pick the single most expensive defect type, at a single point on the line. Do not start by inspecting everything at once.

How to choose the point: the piece sits in a reasonably consistent position when photographed, the conveyor speed is steady, and there is room to mount a camera and a light without rebuilding anything.

Step 2 — Control the lighting before buying an expensive camera

This is the most overlooked step and the number-one cause of technical failure.

Light in a plant changes all day. Window light in the morning is not the same as in the afternoon. Lamps degrade over time. The shadow of someone walking past changes the image. A system trained under one lighting condition degrades immediately when the light changes.

Investing in controlled lighting pays off more than a camera upgrade almost every time, and costs far less. Build an enclosure around the imaging point to shut out ambient light, use LEDs with steady output, and measure the light level periodically.

Step 3 — Run it alongside the existing method

Do not let the system decide anything at this stage. Have it watch and record, in parallel with the existing inspection.

Every week, compare: the pieces a person called defective, did the system see them? The pieces the system flagged, does the person agree? Where they disagree, why?

This period should run at least four to six weeks, and it is the most valuable part of the project, because every disagreement is information that improves the system — and it happens with no risk to production at all.

A by-product that appears every time: the comparison reveals that individual inspectors are applying different standards, which is useful information even if the AI project goes no further.

Step 4 — Let it decide only where it is confident

Once the comparison is stable, let the system start deciding, but only where confidence is above the threshold. Everything else goes to a person.

Set the threshold strictly and relax it gradually. That is better than setting it loosely and having to repair trust you have already lost.

Step 5 — Then expand

Add one defect type at a time, or one inspection point at a time, and go back through step 3 for each addition, every time.

The by-product that often matters more than the original goal

In nearly every successful project, the thing the customer talks about most a year later is not the number of defects caught.

When every piece has been photographed with a timestamp and a lot number, arguments about claims end.

A customer calls to say this lot has a problem. Instead of a dispute about whether the defect happened at the plant or in transit, it becomes a matter of opening the image and seeing what that piece looked like when it left the line.

And once enough data accumulates, patterns appear that were never visible before — one defect type occurring more often on the night shift, or rising every time the raw material lot changes. Sampling can never show this, because sampling keeps nothing.

When to start, and when not to

A good fit if: you produce repeatedly and in volume · defects are visible to the eye · you have customer claims whose cause you could not determine · you want to increase output without adding QC staff in proportion.

Not yet a fit if: you make one-off bespoke items · the defects are not visible, such as internal strength or chemical composition, which need entirely different instruments · volume is low enough that people already inspect every piece · or nobody in the plant is available to own the system.

That last one matters more than it sounds, and it is why we advise some customers to postpone.

In short

Camera-based quality projects do not fail because the AI was not clever enough. They fail because they started at the wrong point, had no images to learn from, could not control the lighting, let the system decide too early, and had nobody in the plant owning them.

All five are fixable by getting the order right, and the most valuable first step — collecting images of defects — can begin with resources already available, while still accounting for staff time, equipment, storage, and data-governance costs.


Want to talk through the process before deciding? See the details at AI for manufacturing, or contact us to request a scope assessment.

Read next: Managed AI operations after go-live

Send sample documents for assessment All articles

Read next