Oct 2022
Perception and Conception
Give a machine eyes and it can leave its cage.

Why distant perception matters
Vision is essential for an organism to understand the world beyond touch. More than 50 per cent of the cortex is devoted to processing visual information, and machine vision is just as hungry for computing power.
Vision is very challenging, but adds enormous capability. Machine vision enables machines to perceive their surroundings, to recognise individuals and objects, to understand context, to discern the attributes of things, and to freely navigate environments. Without vision, machines could not know where to find objects to pick them off a conveyor or out of a box. Machines could not perform quality control on stock to check for damage or missing pieces. Nor could they recognise a catastrophic error, or notify a human colleague.
Advanced machine vision allows for greatly improved mobility and independence. Traditional industrial robots live in a literal cage and cannot easily be repositioned, let alone reposition themselves. Modern robots take themselves wherever they anticipate the greatest need, with minimal oversight or correction necessary from a human guide.
Virtual worlds are growing rapidly in sophistication, and the killer app of the metaverse is not entertainment — it is teaching robots. Humans and machines can work together in a virtual sandbox, humans teaching machines how to act in a certain situation, location, or context. By learning in a virtual environment, we can quickly and cheaply demonstrate a very wide range of potential scenarios. Once machines have gained that experience, they can put it to work in the real world immediately, folding laundry, or recognising uniquely deformed empty drinks cans as rubbish.
The wide variety of affordable and powerful GPUs (graphics cards) has transformed machine vision, due to their speed, parallel processing on thousands of cores, and relative compactness and energy efficiency. Finally we have the raw computational power to enable machines to explore the world in real time, at a high frame rate.
Deep learning techniques have been transformative for machine vision these past ten years, particularly Convolutional Neural Networks (CNNs), which are well suited to vision tasks. However, in the past few years we have seen the emergence of a new generation of machine intelligence techniques, such as Transformers, which are capable of doing lots of different tasks in one model, unlike the deep but narrow focus of earlier networks such as CNNs. Transformers are now eating up even specialist domains, and doing a better job of it.
We can expect that the future of machine vision will be a blend of onboard and remote (cloud) intelligence. Onboard will be used for time-sensitive purposes, and remote intelligence will aid recognition of context and decision making, making sense of situation updates and sending back advice.
Machine vision has advanced enormously, yet significant limitations remain in robustness, context, bias, and performance outside familiar conditions. Productising these developments still requires careful safety and ethical constraints, especially where greater mobility and autonomy can create greater liability.
The Cambrian Explosion, some 540 million years ago, coincided with the evolution of the first eyes, which let primitive animals understand their environment at a distance. It seems that an ability to make sense of multiple modalities of data in physical space will create a similar rapid expansion in capability in machines.
Correspondence
Or send Nell a private note (only Nell and the editorial team see it).
← All essays