Skip to main content
FRAMOS Logo
3 min read

When Detection Isn’t Enough: Object-centric Perception for Robotic Manipulation

FRAMOS

FRAMOS

11. August 2026

When Detection Isn’t Enough: Object-centric Perception for Robotic Manipulation

In a benchmark, perception ends with a label and a bounding box. On a robot, that is where the problem starts. To manipulate an object, a system needs to know not just what something is, but where it is in space, how it is oriented, and whether it is partially hidden – and it needs to keep that estimate stable while the robot, the object, and the scene all move.

At ImagingNext 2026, Maximilian Durner, Research Group Leader in the Perception and Cognition Department at the German Aerospace Center (DLR), presents “Object-centric Perception for Robotic Manipulation” – a session on what it takes for robotic perception to hold up outside structured environments.

Maximilian Durner

About Maximilian Durner

Maximilian Durner is a Research Fellow and Research Group Leader in the Perception and Cognition Department at the Institute of Robotics and Mechatronics of the German Aerospace Center (DLR), where he has been working since 2016. He received his Ph.D. from the Technical University of Munich (TUM). His research group develops methods for robust robotic perception and semantic scene understanding for complex, real-world environments.

Why this matters now

Robotic manipulation is moving out of the structured cell and into environments that were never arranged for robots – logistics, service, and lab settings with open object sets, clutter, occlusion, and changing conditions. In those environments, perception is usually the limiting factor. The object-centric approach treats the object, rather than the image frame, as the unit of understanding: detection, recognition of novel objects, pose estimation, and tracking feed one coherent representation that a manipulation planner can actually act on.

Durner’s group at DLR’s Institute of Robotics and Mechatronics develops methods for robust robotic perception and semantic scene understanding, enabling robots to operate reliably in complex, real-world environments. His research focuses on object-centric vision systems for mobile manipulation, with an emphasis on robustness, continual adaptation, and long-term autonomy – integrating physical priors, world knowledge, and multimodal sensory information to improve capabilities such as object detection, recognition of novel objects, pose estimation, and tracking.

What you’ll take away

  • What “object-centric” perception means in practice, and how it differs from frame-level detection pipelines.
  • How physical priors, world knowledge, and multimodal sensing are combined to make detection, pose estimation, and tracking robust.
  • Where perception still breaks in real-world manipulation – and what robustness and continual adaptation demand from the vision system.

When Detection Isn’t Enough: Object-centric Perception for Robotic Manipulation

Maximilian’s session is one of the talks at ImagingNext 2026 – two days on end-to-end Vision AI systems, edge deployment, and honest engineering exchange. October 14–15, smartvillage Bogenhausen, Munich