From “can see” to “can use”
The biggest question around AR glasses in industry has for years been: what problem do they actually solve? Over the past 18 months the answer has sharpened. Once multimodal models can interpret video and speech in real time, a first-person terminal stops being just a display and becomes the most natural interface between a model and the physical world.
Frontline hands are usually occupied — holding instruments during inspection, tools during maintenance, controlling a scene during enforcement. Any device that demands a free hand gets abandoned. Voice plus first-person capture is one of the few interaction models that adds no operational burden.
Three shifts underway
- Recognition moves forward: video used to be reviewed by humans after the fact; now face, plate and hazard detection happen on-device or at the edge, with anomalies auto-tagged.
- Knowledge moves forward: manuals, procedures and historical tickets move into a knowledge base, and staff ask questions in natural language with answers read back aloud.
- Expertise moves forward: remote experts annotate the first-person view directly, collapsing the cost of a business trip into the cost of a video call.
The deciding factor is not the spec sheet
Weight, battery life and field of view matter, but three other things decide whether a deployment survives: whether data can stay inside the customer’s intranet, whether business workflows can be imported in a standardized way, and whether frontline staff will actually wear it every day.
That is why JoHan keeps repeating “scenario first, delivery above all” on projects — a spec sheet never solved a customer’s problem; a working closed loop does.
This column is industry commentary and technical opinion, and does not constitute investment advice.