Evaluate an AI AOI system by reproducing its claims on your own boards, not by comparing spec sheets. Bring six categories of board (your hardest high-mix job, same-color components, marked inductors, a job with no CAD data, a board with known real defects, and your current worst false-call offender), then measure five things: programming time, false-call rate, escape rate on known defects, operator ramp-up, and run-to-run repeatability. Score against your existing line as the baseline, not against the vendor's numbers.
Every AOI vendor publishes accuracy figures, ours included. Those figures are worth exactly as much as your ability to reproduce them on your product. This article is the protocol we think buyers should run: what to bring, what to measure, how long to run it, and what a passing result looks like. It works on any vendor's machine, and we would rather you run it on ours than take our word for anything.
Why can't you trust AOI accuracy numbers?
AOI accuracy numbers are rarely comparable because the denominators differ. Detection rates get quoted per component, per board, or per defect class, and those three bases can differ by orders of magnitude on the same machine. False-call rates have the same problem. Published figures also omit which boards, which defect set, and which threshold philosophy produced them. The numbers are not necessarily dishonest. They are just not portable to your line.
A quick survey of public claims shows the spread. Independent and vendor sources in 2025 and 2026 variously report 30 to 40 percent fewer false calls, 70 to 85 percent fewer false rejects, and 98 to 99 percent detection on critical solder joints, depending on the source (Overview.ai 2025; Boolean & Beyond 2026; Jidoka 2026). Our own brochure states up to 80 percent fewer false calls and 98 percent or higher detection accuracy (DaoAI-reported). These ranges do not contradict each other so much as they describe different experiments that nobody documented in enough detail to repeat.
Three questions dissolve most of the confusion, and you can ask them of any number, including ours:
- Per what? Per component inspected, per board, or per flagged event. Insist on the denominator.
- On which boards? A demo board designed to look good, or a high-mix production job with second-source parts.
- At which threshold setting? Detection and false calls trade against each other. A number quoted without its counterpart tells you nothing.
None of this means testing is hopeless. It means the test has to happen on your floor.
What boards should you bring to an AOI demo?
Bring six categories of board, each targeting a specific capability: your hardest high-mix job, a board with same-color components, one with marked chip inductors or crystal oscillators, a job you have no CAD data for, a board with known real defects seeded or documented, and the board that generates your worst false-call rate today. Vendors will offer their own demo boards. Those tell you what the machine does well, which is not the question you are trying to answer.
| Board to bring | What it tests | What a good result looks like |
|---|---|---|
| Hardest high-mix job | Programming time under real complexity | Setup completes without vendor intervention |
| Same-color components | Whether detection survives when the body matches the laminate | Parts located reliably, no cluster of misses or flags |
| Marked inductors or crystals | Discrimination between surface markings and defects | Markings not reported as scratches or damage |
| Job with no CAD data | Whether CAD is genuinely optional or quietly required | A working program is produced anyway |
| Board with known defects | Actual escape behavior | Every known defect is caught and correctly classed |
| Current worst false-call board | Whether the pain you have today actually improves | Measurably fewer flags than your existing system |
The known-defect board deserves a note. You need documented, real defects, not simulated ones: boards pulled from your own scrap with the defect list confirmed by your verify station against IPC-A-610 criteria. Keep the list to yourself during the demo. A system evaluated against defects the vendor knows about in advance is not being evaluated.
Also bring the boring board. One straightforward, high-yield job establishes your false-call floor. If a system generates noise on the easy job, nothing it does on the hard ones will save you.
What should you actually measure?
Measure five things: programming time (stopwatch from receiving the board to first inspection), false-call rate (flags divided by components inspected, one consistent denominator), escape rate (known defects found divided by known defects present), operator ramp-up (whether a non-engineer completes a changeover unaided), and repeatability (the same board run three times, comparing results). Record the numbers as you go. Impressions from a demo room fade within a day.
The five, with the arithmetic spelled out:
- Programming time. Start the clock when the board reaches the machine, stop it when the first inspection begins. Include everything: data import, library work, region drawing, threshold setting. Vendor claims about setup speed usually exclude the parts that take the longest. For context on what the steps actually are, see how CAD-free programming works.
- False-call rate. False calls divided by components inspected, expressed in PPM or percent. Pick one denominator and use it for every system you evaluate, including your incumbent. This single discipline makes more comparisons valid than any other step in the protocol.
- Escape rate. Known defects detected divided by known defects present, from your documented defect board. This is the number that protects you. A system that halves false calls while missing one real defect class has not helped you.
- Operator ramp-up. Have a line operator, not a process engineer and not the vendor's applications engineer, run a changeover start to finish. Time it. Note every moment they need help. If the workflow only works in expert hands, your staffing plan just changed.
- Repeatability. Run the same board three times without touching settings. Compare the flags. Variation between identical runs is measurement-system noise, and it caps how much any threshold tuning can accomplish. This is the AOI version of a Gage R&R check, and it is worth doing on your incumbent machine too, because the baseline may surprise you.
Write results on a single sheet during the session. Vendor A versus Vendor B versus your current line, same five rows.
How long should a pilot run?
A two-hour demo can measure programming time, obvious false calls, escape behavior on known defects, and operator ramp-up. It cannot measure whether false calls decline over time, how the system handles a component lot change, or the true escape rate at production volume. Those need two to four weeks on the line. Treat the demo as a screen and the pilot as the decision.
What each stage can actually tell you:
Two-hour demo, honest scope. Setup time is real and measurable here. So is the same-color and marking behavior, because those are single-image problems. Operator ramp-up shows up immediately. Escape behavior on your known-defect board is valid as far as it goes, which is a sample of one board.
Two to four weeks on the line, what only this reveals. Whether the false-call rate trends down as operators feed back corrections, which is the mechanism that separates learning systems from static ones (the mechanics are here). Whether a new component lot or a second-source part breaks anything. Whether the escape rate holds at volume rather than on a single sample. Whether the maintenance burden is what the vendor described.
Two warnings about conclusions people draw too early. A demo cannot validate a learning claim, because learning needs production feedback and time. And a demo running on a vendor-prepared machine says nothing about what your team can maintain three months in. Ask what the machine looked like before you walked in.
What does a passing result look like?
Set the pass mark relative to your own line, not to an absolute figure. Baseline your current programming time, false-call rate, escape rate, and re-inspection headcount, then decide in advance how much improvement justifies the purchase and the disruption. Writing the threshold down before the demo is what keeps the decision honest, because every system looks impressive when the vendor is driving.
A scoring sheet worth filling in before anyone visits:
| Metric | Your current baseline | Minimum to justify a change | Demo result | Pilot result |
|---|---|---|---|---|
| Programming time per job | ___ | ___ | ||
| False-call rate (same denominator) | ___ | ___ | ||
| Escape rate on known defects | ___ | ___ (usually: no regression) | ||
| Operator can run changeover | yes / no | ___ | ||
| Run-to-run repeatability | ___ | ___ |
Two rules for filling it in. First, the escape row is a gate, not a trade: a system that improves everything else while letting real defects through fails, regardless of the other columns. Second, decide the middle column before the demo. A threshold set afterward is a rationalization.
If you are still narrowing the field on machine type and configuration rather than validating claims, the selection guide covers that stage, including the seven questions worth asking every vendor.
Frequently Asked Questions
What if a vendor will not let me bring my own boards?
Treat it as an answer. Confidentiality constraints are real and can be handled with an NDA, but a vendor unwilling to run your product under any arrangement is telling you their numbers live on their boards. Every serious evaluation runs on the buyer's product.
How many boards do I need for the results to mean anything?
For a demo, coverage across the six categories matters more than volume. For a pilot, you want enough production boards that a false-call rate is stable rather than a fluke, which in practice means running the line normally for two to four weeks rather than counting to a specific number.
Can the vendor's engineer tune the system during a pilot?
Yes, and let them. Just record every adjustment they make. That log is your maintenance forecast: everything the applications engineer does during the pilot is work someone in your building will own afterward. If the log is long, ask who does it in month six.
Will an AI system that performs well in a pilot degrade later in production?
The failure mode to ask about is drift: what happens when component lots, board revisions, or lighting change after the model was trained. Systems that learn from operator feedback are designed to absorb that, and DaoAI reports false-call rates declining with use rather than rising (DaoAI-reported). Ask any vendor how the model updates, who triggers the update, and what happens on a board revision.
Run the Protocol on Us
Run this protocol on whatever machines are on your shortlist. If ours is one of them, bring the six boards, keep your defect list to yourself, and hold us to the same scoring sheet as everyone else.
See the P Series Download the Technical Brochure