SMART SENSING insights
Embodied AI brings artificial intelligence into the physical world.

Embodied AI: How to Teach AI to Engage with the World

Are we already at a point where AI is beginning to understand and shape the physical world? Researchers exploring Embodied AI – aka “intelligence with a body” – are starting to push exactly in that direction. We talked to Thomas Wittenberg, Chief Scientist & Research Manager at the Fraunhofer Institute for Integrated Circuits IIS and Group Leader for Visual Healthcare Computing at Friedrich-Alexander-Universität Erlangen-Nürnberg, about this current research avenue. With a track record of +30 years in biomedical engineering and medical informatics, he has a unique take on the potential of Embodied AI.

To start off: what exactly is Embodied AI?

Thomas Wittenberg: Embodied AI is still a relatively new idea. In simple terms, it means taking Artificial Intelligence out of the purely digital world and giving it some kind of physical presence. Up to now, AI has mostly lived in devices like smartphones, laptops, or big servers. What it didn’t have was the ability to actually interact pro-actively with the world around it.

Intelligence needs a body if it’s going to interact with the real world.

A comparison may help: Systems like ChatGPT are incredibly capable, but they’re still basically stuck inside a computer. They interact via prompts and text – and that’s it. Humans, on the other hand, use speech, facial expressions, gestures, notice subtle signals, and physically interact with their surroundings. Embodied AI is moving in that direction. The AI learns to see, hear, and interpret its surroundings. And it gets actuators – for example robotic arms or speech capabilities – so it can actually do something. In a simple scenario, it hands over a glass of water when asked; in a more advanced one, it notices that someone’s voice sounds hoarse and responds proactively – for example, by offering a drink.

But haven’t robots done this for years?

Yes, but the key difference is how you think about it. Traditional robots are basically a shell that engineers try to fill with some type of intelligence. Their capabilities are largely tied to their physical appearance – for example a robotic arm. Embodied AI flips the concept: The intelligence comes first, and then it gets a body – or at least parts of one, if we take the human body as a reference.

Training a model like ChatGPT already involves huge effort und resources. What extra challenges does Embodied AI add? And what can Fraunhofer IIS contribute?

At Fraunhofer IIS, we don’t build robots. But we do have several important areas of expertise that are essential for Embodied AI systems.

First: neuromorphic hardware.

Based on sensory input such as video or audio streams, robots have to react to incoming events in real time. That means its “intelligence” has to run on the device itself. The system can’t wait several seconds for the sensory data to be sent to a cloud server and wait for a response – that latency is far too high for physical interaction.

That’s where neuromorphic hardware comes in. These are very efficient, low‑latency chips that enable AI models to be executed directly on the robot. Large industrial robots can simply be extended and enhanced by multiple GPUs, but smaller and more agile robots need dedicated fast and lightweight hardware, allowing humanoid robots to make complex decisions on the fly. That’s the kind of hardware platform we are working on at Fraunhofer IIS.

Thomas Wittenberg, Chief Scientist & Research Manager at Fraunhofer IIS (©Fraunhofer IIS)

Second: sensor technology.

Small and intelligent sensor systems have been a core topic at Fraunhofer IIS for more than 25 years – whether it’s environmental sensors, contact‑free sensors, sensors worn on the body, IMUs (intertial measurement units), or distance sensing. Embodied AI systems – such as humanoid robots – must constantly perceive and interpret their surroundings. They need to know where they are, whether there are obstacles to avoid, or how the person in front of them is doing. This requires a wide range of different sensors as well as data fusion. A robot must know where it is, who or what is surrounding it, whether there are obstacles, or how the person in front of it is doing. Without good sensor data, embodied systems remain “dumb”.

Third: bridging sensor data and AI.

Even though one focus of Fraunhofer IIS is designing and building large foundation models, our expertise lies in connecting the pieces: getting the sensor data into AI models efficiently, enabling AI to run on small, dedicated hardware platforms in real time, and eventually turning their output into physical actions. This “glue layer” between perception and action has been a research topic for the past decades.

Embodied AI can become a living repository of professional expertise.

With your background in biomedical engineering, can you give us a concrete example where Embodied AI could make a difference?

Imagine we were able to capture and preserve the tacit knowledge that experienced healthcare professionals – for instance nurses or surgeons – acquire over decades of practice. Much of this expertise cannot simply be written down in a textbook. It is implicitly embedded in movements, routines, and intuitive decisions developed through years of patient care.

This is precisely the aim of our current research project, ROBO.KIWI. Using body-worn sensors, we capture how experienced caregivers move, interact with patients, and perform complex tasks. Combined with AI, this knowledge is then curated and can be transferred to humanoid robots. Within nursing education, such a robot could take on a dual role: as a caregiver, it demonstrates techniques and movement sequences; as a care recipient, it can simulate realistic patient responses and physical resistance. In this sense, Embodied AI can become a living repository of professional expertise.

That sounds particularly relevant in light of demographic change and the growing shortage of nursing professionals. You mentioned body-worn sensor networks, which links directly to research on motion and gait analysis and the Wireless Body Area Network (WBAN) maphera®. Are these technologies building blocks for Embodied AI?

Yes, definitely. You can look at motion and movement analysis from two sides.

One area relates to diagnosis, therapy monitoring, and rehabilitation support. Today, such assessments often rely on high-end motion and gait analysis labs equipped with cameras or motion caption suits, such as we operate in our Center for Sensor Technology and Digital Medicine. A wearable sensor network like maphera®, however, could do the same job in everyday settings. With a combination of synchronized vital sensors and IMUs, it becomes possible to monitor whether patients are walking correctly after surgery or performing rehab exercises the right way.

Movement is data – and it’s incredibly valuable.

The other side is so-called expert knowledge, as mentioned in connection with our ROBO.KIWI project. The way nurses, physiotherapists, or surgeons move and interact with patients or objects contains a lot of subtle, valuable information about tacit “motion knowledge”. Using adequate motion capture technologies – from optical systems to body-worn sensor networks – it is possible to record, preserve, curate, and analyze this knowledge. And once a sufficiently rich collection of curated motion patterns and clinical skills is available, it can be used to train exoskeletons, teach humanoid robots, or create novel training and educational tools. In the end, it’s about capturing human motion expertise that would otherwise be lost.

In my opinion, both aspects are highly important steps toward Embodied AI and humanoid robotics.

If we solve the issue of motion and context awareness, Embodied AI takes a leap forward.

What are the main technical hurdles right now? What are the next steps in your research toward Embodied AI?

Mobile Embodied AI systems such as humanoid robots need both, motion and context awareness. As mentioned, they need to know where they are, where they are moving to and how fast, and how they are oriented in space. That’s where kinesthetics – the understanding of movement intelligence – comes in. Currently, we are able to capture these parameters in our motion lab, using optical motion‑capture systems and motion capture suits with integrated IMUs. But these approaches have limitations: High-end motion capture systems are mainly stationary and confined to specialized lab environments. Wearable sensor networks allow us to capture natural movements more easily and in a much wider range of situations.

This mobility offers a major advantage, especially when working with less cooperative subjects such as patients. Wearable sensor technology allows us to provide Embodied AI systems with the real-world motion data they need to understand and interact with their environment.

Thank you for the insights, Thomas. We’ll be watching closely where your research takes you next.

Image copyright (cover image): Fraunhofer IIS / Paul Pulkert

Grit Nickel

Grit Nickel

Grit is a content writer at Fraunhofer IIS and a technology communication specialist. She has 6+ years of experience in research and holds a PhD in German linguistics.

Add comment

All Categories