Oct 2022
Perception and Conception
Give a machine eyes and it starts to understand the world, not just describe it.
Why distant perception matters
Vision is essential for an organism to understand the world beyond touch. More than 50 percent of the cortex is devoted to processing visual information -- and the ratio in robots and other artificial intelligences is not so different.
Vision is very challenging, but adds enormous capability. Machine vision enables machines to perceive their surroundings, to recognise individuals and objects, to understand context, to discern the attributes of things, and to freely navigate environments. Without vision, machines could not know where to find objects to pick them off a conveyor or out of a box. Machines could not perform quality control on stock to check for damage or missing pieces. Nor could they recognise a catastrophic error, or notify a human colleague.
Advanced machine vision allows for greatly improved mobility and independence. Traditional industrial robots live in a literal cage and cannot easily be repositioned (let alone repositioning themselves). Modern robots take themselves wherever they anticipate the greatest need, with minimal oversight or correction necessary from a human guide.
Virtual worlds are growing rapidly in sophistication, and the killer app of the metaverse is not entertainment -- it is teaching robots. Humans and machines can work together in a virtual sandbox, humans teaching machines how to act in a certain situation, location, or context. By learning in a virtual environment, we can quickly and cheaply demonstrate a very wide range of potential scenarios. Having gained experience, that learning can be immediately put to work in the real world, enabling machines to fold laundry, or recognise uniquely deformed empty drinks cans as trash.
The wide variety of affordable and powerful GPUs (graphics cards) has transformed machine vision, due to their speed, parallel processing on thousands of cores, and relative compactness and energy efficiency. Finally we have the raw computational power to enable machines to explore the world in real time, at a high frame rate.
Deep learning techniques have been transformative for machine vision these past ten years, particularly Convolutional Neural Networks, which are well suited to vision tasks. However, in the past few years we have seen the emergence of a new generation of machine intelligence techniques, such as Transformers, which are capable of doing lots of different tasks in one model, unlike the deep but narrow focus of deep learning. Transformers are now eating up even specialist deep learning domains, and doing a better job of it.
We can expect that the future of machine vision will be a blend of onboard and remote (cloud) intelligence. Onboard will be used for time-sensitive purposes, and remote intelligence will aid recognition of context and decision making, making sense of situation updates and sending back advice.
Machine vision has advanced enormously, yet significant limitations remain in robustness, context, bias, and performance outside familiar conditions. Productising these developments still requires careful safety and ethical constraints, especially where greater mobility and autonomy can create greater liability.
The Cambrian Explosion 530-545 million years ago occurred when eyes first evolved, enabling primitive animals to understand their environment at a distance. It seems that an ability to make sense of multiple modalities of data in physical space will create a similar rapid expansion in capability in machines.
Correspondence