Machine vision in manufacturing is inspection: a camera and software decide whether a part is acceptable and hand that decision to an automation system. This collection is organised around a single problem — the published evidence about how well it works is much weaker than the volume of writing about it suggests, and the strongest sources are often the ones that say the least.
Start with what the systems do, because it decides how to read every number that follows. The industry's own shorthand is GIGI — Guidance, Identification, Gauging, Inspection — which a vendor breaks into six application classes: defect detection, object detection and counting, measuring and gauging, locating and guiding, barcode reading, and character recognition and verification. These are not commensurable. A barcode has a ground truth that is checkable character by character; whether a faint mark on a metal surface is a defect can be hard for an expert to adjudicate. An accuracy figure from one class says nothing about another.
Then the part that decides whether any of it works. Cognex states that poor lighting is the most common cause of poor machine vision performance, and that sophisticated cameras and software cannot make up for it. The design requirement is that illumination maximise contrast on the feature of interest and stay consistent against normal variation in parts and their arrangement — which is the whole failure mode written as a specification, since a system is only as stable as the lighting it was tuned under.
The industry does have a public standard, and what it covers is instructive. EMVA 1288, at Release 4.0 since June 2021, defines how to measure and present what a camera does: quantum efficiency, temporal dark noise, dynamic range, spatial non-uniformity, defect pixels. It standardises the sensor. No source read for this collection describes an equivalent standard for reporting how often a deployed system rejects a good part or passes a bad one — a statement about what was read here, not a claim that no such framework exists.
On deployments, the honest summary is that scale and method are public, performance is not, and the sample is very small. One manufacturer-authored machine vision record was retrievable this session: BMW, which has run AI image recognition in series production since 2018, built from around 100 photographs per feature taken by employees on a mobile camera. It publishes no false-reject rate, no escape rate, no test-set description and no line speed, and the one figure it does give — that reliability reaches 100% after a test run — defines no metric or test set. That absence is the finding. One record is not a survey, and this collection does not generalise beyond it.
What measurement does exist points one way. On 2,042 real metal-box images in an unconstrained industrial environment, the best method in a peer-reviewed study achieved 10.6% false positives and 5.41% false negatives — roughly one good part in ten pulled for review, roughly one bad part in twenty getting through. A 2025 benchmark study is blunter about why published figures mislead: across nine datasets, eleven models and seven metrics, models reaching 99.9% image-level AUROC on the field's standard academic dataset degrade significantly on real production data, and its own corrective benchmark excludes that dataset entirely.
Two pairs of sources are held side by side here without being forced into conflict. BMW's 2019 release states that pseudo-defects — false alarms from dust or oil — no longer occur in one press-shop application; practitioners in 2026 call pseudo-defects the most common reason these deployments fail across the field. A solved instance and a general failure rate are not contradictory, and neither source speaks to the other's scope. Likewise the benchmark study argues a missed defect is the costlier event while the practitioners argue repeated false alarms are what end deployments — a question about the cost of one error and a question about the frequency of many, both of which can hold at once.
Where inspection breaks down is answered here mainly from the research side: the benchmark study treats robustness under distribution shift as one of its open experiments, and practitioners describe systems reacting to lighting shifts, reflections and material batch changes. One adjacent record is included with its boundary stated on the page — Audi's spot-weld system, which is NOT established as machine vision, since its release identifies no camera or image sensor and the method it replaced was ultrasound. Audi states that moving that system between Volkswagen Group plants required retraining for each site's weld settings. That is suggestive about AI inspection generally and is not evidence about vision.