Topicsmachine-vision

Machine Vision: How Automated Inspection Is Built and Judged

A sourced reference collection on automated visual inspection in manufacturing — what the systems are asked to do, what they physically require, how their results should be judged, and why the measured error rates look nothing like the accuracy claims.

Assertions
15
Sources consulted
12
Read in full
10/12
Cited as evidence
10

12 sources sit behind this page — including any that arrive with a concept this page shares with another collection. 10 were retrieved and read in full, and only those can back an assertion. 1 could not be retrieved, and 1 were surfaced and deliberately set aside. Every one of them is named in the register below, with the reason in view. How we source this.

Background

Our own synthesis, written to orient you — not evidence. Every factual statement here is asserted and sourced further down this page.

Machine vision in manufacturing is inspection: a camera and software decide whether a part is acceptable and hand that decision to an automation system. This collection is organised around a single problem — the published evidence about how well it works is much weaker than the volume of writing about it suggests, and the strongest sources are often the ones that say the least.

Start with what the systems do, because it decides how to read every number that follows. The industry's own shorthand is GIGI — Guidance, Identification, Gauging, Inspection — which a vendor breaks into six application classes: defect detection, object detection and counting, measuring and gauging, locating and guiding, barcode reading, and character recognition and verification. These are not commensurable. A barcode has a ground truth that is checkable character by character; whether a faint mark on a metal surface is a defect can be hard for an expert to adjudicate. An accuracy figure from one class says nothing about another.

Then the part that decides whether any of it works. Cognex states that poor lighting is the most common cause of poor machine vision performance, and that sophisticated cameras and software cannot make up for it. The design requirement is that illumination maximise contrast on the feature of interest and stay consistent against normal variation in parts and their arrangement — which is the whole failure mode written as a specification, since a system is only as stable as the lighting it was tuned under.

The industry does have a public standard, and what it covers is instructive. EMVA 1288, at Release 4.0 since June 2021, defines how to measure and present what a camera does: quantum efficiency, temporal dark noise, dynamic range, spatial non-uniformity, defect pixels. It standardises the sensor. No source read for this collection describes an equivalent standard for reporting how often a deployed system rejects a good part or passes a bad one — a statement about what was read here, not a claim that no such framework exists.

On deployments, the honest summary is that scale and method are public, performance is not, and the sample is very small. One manufacturer-authored machine vision record was retrievable this session: BMW, which has run AI image recognition in series production since 2018, built from around 100 photographs per feature taken by employees on a mobile camera. It publishes no false-reject rate, no escape rate, no test-set description and no line speed, and the one figure it does give — that reliability reaches 100% after a test run — defines no metric or test set. That absence is the finding. One record is not a survey, and this collection does not generalise beyond it.

What measurement does exist points one way. On 2,042 real metal-box images in an unconstrained industrial environment, the best method in a peer-reviewed study achieved 10.6% false positives and 5.41% false negatives — roughly one good part in ten pulled for review, roughly one bad part in twenty getting through. A 2025 benchmark study is blunter about why published figures mislead: across nine datasets, eleven models and seven metrics, models reaching 99.9% image-level AUROC on the field's standard academic dataset degrade significantly on real production data, and its own corrective benchmark excludes that dataset entirely.

Two pairs of sources are held side by side here without being forced into conflict. BMW's 2019 release states that pseudo-defects — false alarms from dust or oil — no longer occur in one press-shop application; practitioners in 2026 call pseudo-defects the most common reason these deployments fail across the field. A solved instance and a general failure rate are not contradictory, and neither source speaks to the other's scope. Likewise the benchmark study argues a missed defect is the costlier event while the practitioners argue repeated false alarms are what end deployments — a question about the cost of one error and a question about the frequency of many, both of which can hold at once.

Where inspection breaks down is answered here mainly from the research side: the benchmark study treats robustness under distribution shift as one of its open experiments, and practitioners describe systems reacting to lighting shifts, reflections and material batch changes. One adjacent record is included with its boundary stated on the page — Audi's spot-weld system, which is NOT established as machine vision, since its release identifies no camera or image sensor and the method it replaced was ultrasound. Audi states that moving that system between Volkswagen Group plants required retraining for each site's weld settings. That is suggestive about AI inspection generally and is not evidence about vision.

Figures

Every number below is asserted and sourced elsewhere on this page.

What the errors actually look like

Best-performing method on 2,042 real metal-box images in an unconstrained industrial setting. About one good part in ten is pulled for review; about one bad part in twenty gets through.

Error rate on 2,042 real images, per cent

10.6%

False positives

5.41%

False negatives

Peer-reviewed study of defect detection on metal boxes captured in an unconstrained industrial environment. Measured rates for the best method tested on that dataset — not a general figure for machine vision, and not transferable to another part or line.

The corrective benchmark, and what it left out

A 2025 benchmark study assembled nine datasets weighted towards defects produced in real production rather than in a laboratory — and excluded MVTecAD, the field's most-used academic dataset, from the benchmark entirely.

Total 9 datasets

  • Defects produced in real-world settings6
  • Defects produced in laboratory settings3

Composition of the nine datasets listed in Table 2 of a non-peer-reviewed preprint (v1, 30 March 2025), counted from the paper's own lab/real-world type column. The count is of datasets, not of images or of results.

Concepts

The vocabulary this subject is built from, and what we can show about each.

EMVA 1288 Camera Characterisation

other

EMVA 1288 gives a unified method to measure and present camera specifications for machine vision — quantum efficiency, system gain, temporal dark noise, saturation capacity, dynamic range, dark current, DSNU and PRNU, defect pixels and signal-to-noise ratio — with Release 4.0 effective June 2021 superseding Release 3.1 of December 2016; it standardises the sensor, and no source read here describes an equivalent standard for the inspection decision.

ReportedSupported by the sources below, not yet editor-reviewed.
3 sources3 retrieved & read

Lighting, Optics and the Imaging Chain

process

Cognex states that poor lighting is the most common cause of poor machine vision performance and that cameras and software cannot compensate for it, that lighting must maximise contrast on features of interest while staying consistent against normal part variation, and that there is no single best setup — naming backlighting, bright field, dark field, diffuse and multispectral techniques plus colour, IR/UV and polarising filters.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Machine Vision

component

On 2,042 real metal-box images captured in an unconstrained industrial environment, the best method in a peer-reviewed study achieved 10.6% false positives and 5.41% false negatives on defect localisation, against 13.02% and 8.6% for a fine-tuned VGG-16 baseline.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

Machine vision is the industrial application of imaging — a combination of sensors, lenses, lighting and software that inspects, measures or identifies a part and hands a decision to an automation system — and its vendors distinguish it from the broader term computer vision, sorting cameras into line-scan, 2D area-scan and 3D categories.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Pseudo-Defect

other

BMW's 2019 release states that pseudo-defects from dust and oil residues at its Dingolfing press shop no longer occur with an AI application given around 100 real images per feature, while 2026 practitioner accounts identify pseudo-defects as the most common cause of AI vision deployment failure across the field — a solved instance and a general failure rate, at different scopes, neither speaking to the other.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Around 77% of AI vision implementations in manufacturing are reported never to advance beyond the pilot phase — a figure described as widely known but rarely examined, published without methodology or an identified source study.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

A pseudo-defect is a false alarm triggered by natural variation such as lighting shifts, reflections or material batch changes; practitioners identify it as the most common cause of AI vision deployment failure, because at production volume even a small false positive rate erodes operator trust until alerts are ignored.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

Rule-Based Versus Learned Inspection

process

Cognex states that rule-based systems suit consistent, high-speed tasks while AI handles variability and eases setup and maintenance; BMW describes its own learned method as employees photographing a component from several angles and marking deviations, with a server computing the network from around 100 images in a training stage that may run overnight.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Visual Anomaly Detection

process

Across nine datasets, eleven state-of-the-art models and seven metrics, a benchmark study found that models reaching 99.9% image-level AUROC on MVTecAD degrade significantly on real-world data; its own benchmark uses six datasets with real-world defects against three with laboratory-produced defects, and excludes MVTecAD.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

The benchmark study argues that image-level AUROC is not accountable for the relative importance of errors, that in production a missed defective part costs significantly more than a false positive, and that test-set-based early stopping, best-epoch reporting and centre-crop augmentation inflate published results — a cost-per-error view that sits alongside, rather than against, the practitioner account that repeated false alarms are what end deployments.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

What Deployment Records Actually Publish

other

Audi's WPS-Analytics spot-weld system is adjacent AI quality-control context and is NOT established as machine vision — the release identifies no camera, optical system or image sensor and the method it replaced was ultrasound — but it states that installing the system at other Volkswagen Group plants required identifying differences in weld settings between sites in order to retrain the model.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

The one manufacturer-authored machine vision deployment record found for this collection — BMW, running AI image recognition in series production since 2018 on named final-inspection and press-shop applications — publishes method and scope but no false-reject rate, escape rate, test-set description, observation period or line speed.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

What Inspection Systems Are Asked To Do

process

The industry shorthand GIGI covers Guidance, Identification, Gauging and Inspection, which Cognex breaks into six application classes — defect detection, object detection and counting, measuring/gauging, locating/guiding/positioning, barcode reading, and OCR/OCV — tasks whose ground truths differ so much that accuracy figures from one class cannot be compared with another.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Timeline

What actually happened, in order, with sources.

Coverage
  • 1 Germany
  • 1 Unattributed

Where this topic’s events took place, as far as our sources establish it. Events with no single location — a standards publication, say — and events we have not yet attributed are both counted as unattributed rather than omitted.

  1. Jun 2021

    The Camera Characterisation Standard Reaches Release 4.0

    otherEurope

    Release 4.0 of EMVA 1288 took effect in June 2021, superseding Release 3.1 of December 2016 and extending camera characterisation beyond linear-response cameras with simple preprocessing.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  2. Jul 15, 2019

    An Automaker Describes AI Inspection in Series Production

    system adoptionGermany

    On 15 July 2019 the BMW Group published an account stating it had used AI applications in series production since 2018, building image recognition from around 100 images per feature photographed by employees on a mobile standard camera, with applications including warning-triangle and wiper-cap checks and model-designation verification at Dingolfing final inspection.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read

Source register

All 12 sources behind this page — what we read, what we tried to read and could not, and what we looked at and set aside, with the reason in view for each. A concept shared with another collection brings its own references with it, so some entries here were surfaced for a neighbouring topic rather than this one.

Cited as evidence
10
Tried, could not read
1
Surfaced, set aside
1
Cited sources 10 distinct links

Original publisher links. Files open on the publisher’s site; we do not host copies. A linked document is not an additional source or an independent verification.

Tried, could not read1

We attempted these and were refused or served nothing. Nothing on this page rests on them; they are published so the gaps are checkable rather than invisible.

Surfaced, set aside1

These came up while researching and were deliberately not used. We do not claim to have read them — each is listed with why it was passed over, so the shape of the survey is visible and not just its conclusions.

Coverage & limits

What this page does and does not claim.

Ninth packet through the generic ingestion pipeline (August 2026), substantially expanded on 14 September 2026 to answer six reader questions the first revision did not address: what separates conventional from learned inspection, what tasks are performed, what a working system requires, what deployment evidence exists, how results should be judged, and where performance breaks down. Ten sources are now directly retrieved against three before. Two existing sources were upgraded: the anomaly-detection preprint was previously abstract-only and its full text was read this session, and Cognex — previously recorded as blocked by HTTP 429 — serves normally to an ordinary browser, so three of its technical pages were read, with the retrieval method recorded. One manufacturer press release was added as vision deployment evidence — BMW, describing image recognition on its own lines — and it remains evidence of what the issuer says rather than independent validation, publishing no measured result. A second, Audi's spot-weld release, is included ONLY as explicitly labelled adjacent AI quality-control context: it identifies no camera, optical system or image sensor, its method predecessor was ultrasound, and no source read establishes an imaging modality for it, so it supports no vision claim here. An earlier draft of this revision treated it as the collection's flagship vision deployment, which was a category error corrected in peer review. Vendor pages are cited only for definitional, structural and requirement claims a vendor is the right authority for; no performance figure is taken from any of them. Two pairs of sources are recorded side by side as complementary rather than contradictory — a scoped pseudo-defect success against a general failure rate, and cost-per-escape against false-alarm frequency — and the collection does not assert an explanation for either pairing, since no source read tests one. Known gaps. The Springer defect-detection review remains behind publisher authentication and was not bypassed, so the measured error rates here still rest on a single 2018 study rather than on a survey of the literature. The final EMVA 1288 Release 4.0 documents are a membership download, so the parameter list here comes from the openly published March 2021 release candidate. The benchmark study's per-model results table was not transcribed because its values flatten ambiguously in extracted text and mis-mapping a cell would manufacture a number. Market-size figures remain deliberately omitted. The most-repeated figure in this collection — that around 77% of AI vision implementations never leave pilot — is still published with its provenance gap stated and no longer carries a chart of its own. Not yet editor-reviewed; every assertion reads as reported.

Source check, 2026-09-17. Numeric-presence checks passed for 15 assertions using available source text, which may be cached. This is not verification of their meaning. What this check does and does not prove →

  • Not editor-reviewed unless labelled. Assertions marked Reported are assembled from the sources shown and have not yet been checked by an editor. Only Primary source and Corroborated mean a human verified them.
  • Disagreements are preserved, not resolved. Where sources conflict, both accounts appear and the assertion is marked Disputed.
  • Retrieval status is disclosed per source. A source we could not open is never counted as evidence for an assertion.

This page is also available as structured data: /api/v1/topics/machine-vision