Header

Search

Student Projects

To apply, please send your CV, your Ms and Bs transcripts by email to all the contacts indicated below the project description. Do not apply on SiROP Since Prof. Davide Scaramuzza is affiliated with ETH, there is no organizational overhead for ETH students. Custom projects are occasionally available. If you would like to do a project with us but could not find an advertized project that suits you, please contact Prof. Davide Scaramuzza directly to ask for a tailored project (sdavide at ifi.uzh.ch).

Upon successful completion of a project in our lab, students may also have the opportunity to get an internship at one of our numerous industrial and academic partners worldwide (e.g., NASA/JPL, University of Pennsylvania, UCLA, MIT, Stanford, ...).

  • Learning Robust Agile Flight via Adaptive Curriculum

    This project focuses on developing robust reinforcement learning controllers for agile drone navigation using adaptive curricula. Commonly, these controllers are trained with a static, pre-defined curriculum. The goal is to develop a dynamic, adaptive curriculum that evolves online based on the agents' performance to increase the robustness of the controllers.

  • Long-Horizon Learning for Agile Autonomous Drone Racing

    This thesis investigates how long-horizon sequence modeling can improve autonomous drone racing. The project focuses on developing learning-based control methods that enable drones to reason over time, adapt to rapidly changing environments, and make high-speed decisions under dynamic constraints.

  • Event‑based Temporal Segmentation & Tracking

    Event cameras are revolutionary sensors that capture pixel-level illumination changes with microsecond latency, providing significant advantages in high-speed and high-dynamic-range scenarios where traditional cameras suffer from motion blur. Recently, large-scale foundational segmentation models have been successfully adapted to the event domain. However, these current approaches remain constrained to per-frame analysis, treating continuous event streams as isolated, static snapshots and ignoring temporal consistency. At the same time, existing event-based methods for moving object segmentation can isolate motion but fail to maintain instance identity over time—they can segment moving pixels, but they cannot "track" specific objects. This project aims to bridge the gap between static foundational segmentation and dynamic motion analysis by developing the first comprehensive tracker for event cameras. The objective is to design a system capable of not only segmenting arbitrary objects but also maintaining their identity consistently across long, high-speed sequences. The student will extend current spatial feature adaptation strategies to support temporal identity, effectively transforming a frame-by-frame instance segmenter into a robust Video Object Segmentation (VOS) tracker. Furthermore, to handle severe object occlusions and rapid, erratic motion, the project will explore sparse temporal memory mechanisms that prevent identity-switching. Finally, to rigorously test the system's reliability, the student will establish a novel benchmark for dense segmentation in extreme edge cases, such as night driving with severe glare and rapid evasive maneuvers.

  • Reinforcement Learning with World Models

    Explore and develop model-based RL algorithms.

  • Vision-based Navigation in Dynamic Environment via Reinforcement Learning

    In this project, we are going to develop a vision-based reinforcement learning policy for drone navigation in dynamic environments. The policy should adapt to two potentially conflicting navigation objectives: maximizing the visibility of a visual object as a perceptual constraint and obstacle avoidance to ensure safe flight.

  • Event Representation Learning for Control with Visual Distractors

    This project develops event-based representation learning methods for control tasks in environments with visual distractors, leveraging sparse, high-temporal-resolution event data to improve robustness and efficiency over traditional frame-based approaches.

  • Rethinking RNNs for Neuromorphic Computing and Event-based Vision

    This thesis develops hardware-optimized recurrent neural network architectures with novel parallelization and kernel-level strategies to efficiently process event-based vision data for real-time neuromorphic and GPU-based applications.

  • Time-continuous Facial Motion Capture Using Event Cameras

    Traditional facial motion capture systems, including marker-based methods and multi-camera rigs, often struggle to capture fine details such as micro-expressions and subtle wrinkles. While learning-based techniques using monocular RGB images have improved tracking fidelity, their temporal resolution remains limited by conventional camera frame rates. Event-based cameras present a compelling solution, offering superior temporal resolution without the cost and complexity of high-speed RGB cameras. This project explores the potential of event-based cameras to enhance facial motion tracking, enabling the precise capture of subtle facial dynamics over time.

  • Vision Language Action models for Drones

    This project explores generative modeling of drone flight paths conditioned on natural language commands and spatial constraints, aiming to produce plausible 3D trajectories for training reinforcement learning policies. We investigate model architectures, data sources, and trajectory extraction methods to ensure generated paths are both physically feasible and stylistically aligned with textual descriptions.

  • Self-Supervised Event-Driven World Models for High-Speed Scene Forecasting

    Predicting how a dynamic scene will evolve is a cornerstone of safe, agile robotic navigation. While conventional self-supervised world models rely on frame-based video prediction, they fail during rapid motions due to motion blur and low sampling rates. Event cameras circumvent these limitations by tracking continuous brightness changes with microsecond latency. This project focuses on developing a Self-Supervised Event-Driven World Model (S-EWM) that learns to forecast future environmental states directly from raw, asynchronous event streams. By predicting future event distributions or synthesized frames without human labels, the model will capture the underlying physics of highly dynamic environments, serving as a powerful representation for downstream robotic perception.

  • Independent Moving Object Segmentation with the Aeveon Sensor

    This project extends the Motion-aware Event Suppression framework by integrating multi-bit event streams and synchronized intensity data from the new iniVation Aeveon sensor to improve real-time object segmentation. It aims to develop a multi-modal neural architecture that disentangles ego-motion from independently moving objects, evaluated through real-world benchmark sequences.

  • Hardware-Aware Mapping and Quantization of Event-Driven GNNs for Edge Vision

    This project focuses on the hardware-aware mapping and quantization of Asynchronous Graph Neural Networks (AGNNs) for edge vision applications. By optimizing DaGR network operations and developing advanced quantization techniques, the research aims to bridge the gap between software simulation and physical execution on the GNN processor, ensuring high performance under strict bit-width constraints.

  • Spiking Neural Networks for Agile Autonomous Drone Racing

    Explore the use of Spiking Neural Networks (SNNs) for agile, autonomous drone racing, focusing on achieving low-latency and energy-efficient flight control. By leveraging event-based processing, the research aims to develop and validate high-performance spiking policies suitable for deployment on neuromorphic hardware.

  • Neural Vision for Celestial Landings (in collaboration with European Space Agency)

    In this project, you will investigate the use of event-based cameras for vision-based landing on celestial bodies such as Mars or the Moon.

  • Learning Suspended Payload Dynamics in the Real World for Agile Maneuvers

    Learn a dynamics model of the suspended payload using real-world data to perform complex maneuvers with unprecedented agility.

  • Fast Manhole Traversal During Ship Inspection Using Drone Racing Techniques

    This project aims to propose a method that utilizes drone racing techniques to traverse manholes in ballast water tanks, without a prior known map of the environment. Depending on the result, the policy may be deployed on a real container ship.

  • Aerobatic Agile Object Grasping and Delivery

    Taking advantage of our high-fidelity simulation of the quadrotor with gripper, train a policy that’s capable of performing aerobatic maneuvers to grasp objects at constrained locations.

  • Limits of Learned SLAM

    This project investigates the failure modes of learned SLAM systems, aiming to identify challenging scenarios, understand why current methods fail, and develop solutions that improve robustness without degrading nominal performance.

  • Event Cameras for Agile Aerial 3D Perception

    Fast drones often move too quickly for conventional cameras, resulting in motion blur and unreliable 3D perception. This project investigates how event cameras, which capture microsecond-level brightness changes, can help drones “see clearly” during aggressive flight. The student will develop learning-based methods that combine standard images, event data, and motion cues to recover sharp visual information and reconstruct the 3D environment.

  • Safe Vision-Based Recovery from Aggressive Quadrotor Flight via Reinforcement Learning

    This project investigates safe quadrotor recovery from aggressive racing conditions using only onboard visual observations. The goal is to learn a policy that can stabilize the drone from high-speed, highly tilted, and obstacle-proximate states without relying on reliable state estimation.

  • Learning-Based Odometry with Multiple IMUs

    This project investigates learning-based camera-less odometry using multiple spatially distributed IMUs to improve motion estimation accuracy and robustness over single-IMU approaches. It will explore how learned models can exploit both sensor redundancy and spatial measurement differences, with magnetometers as a possible extension.

  • Event-based Perception for Autonomous Driving

    This project explores how event-based sensing can enhance perception for autonomous driving systems. It investigates the integration of asynchronous, high-temporal-resolution visual signals with conventional sensing modalities to improve robustness, latency, and reliability under challenging real-world conditions.

  • Merging Vision-Language-Action Models and World Models for Vision-Based Robot Navigation

    What if a robot could follow instructions like "fly through the door and inspect the shelf" — and imagine the outcome of its actions before executing them? This project merges two of the most exciting paradigms in robot learning: Vision-Language-Action (VLA) models, which provide open-world semantic understanding and instruction following, and learned World Models (WM), which predict how the environment evolves and enable foresight.

  • Scaling Optimization Layers for Robust Generalist Vision-Based Drone Navigation

    Learned policies are fast and flexible; model-based optimization is safe and interpretable. Embedded as differentiable optimization layers inside a neural network, both can be trained jointly — but so far only at small scale and in narrow task settings. This project asks: what happens when optimization layers are scaled up across thousands of parallel environments and diverse tasks to train a robust generalist policy for vision-based drone navigation?

  • Vision-Based World Models for Real-Time Robot Control

    This master's project focuses on enabling real-time, vision-based control for quadrotors by distilling large, complex world models into lightweight versions suitable for deployment on resource-constrained platforms. The goal is to achieve fast, efficient inference from camera inputs, supporting agile indoor navigation in previously unseen environments.

  • Vision-Based Tactile Sensor for Humanoid Hands (in collaboration with Soft Robotics Lab)

    Humanoid robots require tactile sensing to achieve robust dexterous manipulation beyond the limits of vision-based perception. This project develops an event-based tactile sensor to provide low-power, high-bandwidth force estimation from material deformation, with the goal of integrating it into a human-scale robotic hand.