Student Projects
To apply, please send your CV, your Ms and Bs transcripts by email to all the contacts indicated below the project description. Do not apply on SiROP . Since Prof. Davide Scaramuzza is affiliated with ETH, there is no organizational overhead for ETH students. Custom projects are occasionally available. If you would like to do a project with us but could not find an advertized project that suits you, please contact Prof. Davide Scaramuzza directly to ask for a tailored project (sdavide at ifi.uzh.ch).
Upon successful completion of a project in our lab, students may also have the opportunity to get an internship at one of our numerous industrial and academic partners worldwide (e.g., NASA/JPL, University of Pennsylvania, UCLA, MIT, Stanford, ...).
-
Learning Rapid UAV Exploration with Foundation Models

Recent research has demonstrated significant success in integrating foundational models with robotic systems. In this project, we aim to investigate how these foundational models can enhance the vision-based navigation of UAVs. The drone will utilize learned semantic relationships from extensive world-scale data to actively explore and navigate through unfamiliar environments. While previous research primarily focused on ground-based robots, our project seeks to explore the potential of integrating foundational models with aerial robots to enhance agility and flexibility.
-
Event Cameras for Agile Drone 3D Perception

Fast drones often move too quickly for conventional cameras, resulting in motion blur and unreliable 3D perception. This project investigates how event cameras, which capture microsecond-level brightness changes, can help drones “see clearly” during aggressive flight. The student will develop learning-based methods that combine standard images, event data, and motion cues to recover sharp visual information and reconstruct the 3D environment.
-
Learning Robust Agile Flight via Adaptive Curriculum

This project focuses on developing robust reinforcement learning controllers for agile drone navigation using adaptive curricula. Commonly, these controllers are trained with a static, pre-defined curriculum. The goal is to develop a dynamic, adaptive curriculum that evolves online based on the agents' performance to increase the robustness of the controllers.
-
Long-Horizon Learning for Agile Autonomous Drone Racing

This thesis investigates how long-horizon sequence modeling can improve autonomous drone racing. The project focuses on developing learning-based control methods that enable drones to reason over time, adapt to rapidly changing environments, and make high-speed decisions under dynamic constraints.
-
Event‑based Temporal Segmentation & Tracking

Event cameras are revolutionary sensors that capture pixel-level illumination changes with microsecond latency, providing significant advantages in high-speed and high-dynamic-range scenarios where traditional cameras suffer from motion blur. Recently, large-scale foundational segmentation models have been successfully adapted to the event domain. However, these current approaches remain constrained to per-frame analysis, treating continuous event streams as isolated, static snapshots and ignoring temporal consistency. At the same time, existing event-based methods for moving object segmentation can isolate motion but fail to maintain instance identity over time—they can segment moving pixels, but they cannot "track" specific objects. This project aims to bridge the gap between static foundational segmentation and dynamic motion analysis by developing the first comprehensive tracker for event cameras. The objective is to design a system capable of not only segmenting arbitrary objects but also maintaining their identity consistently across long, high-speed sequences. The student will extend current spatial feature adaptation strategies to support temporal identity, effectively transforming a frame-by-frame instance segmenter into a robust Video Object Segmentation (VOS) tracker. Furthermore, to handle severe object occlusions and rapid, erratic motion, the project will explore sparse temporal memory mechanisms that prevent identity-switching. Finally, to rigorously test the system's reliability, the student will establish a novel benchmark for dense segmentation in extreme edge cases, such as night driving with severe glare and rapid evasive maneuvers.
-
Reinforcement Learning with World Models

Explore and develop model-based RL algorithms.
-
Vision-based Navigation in Dynamic Environment via Reinforcement Learning

In this project, we are going to develop a vision-based reinforcement learning policy for drone navigation in dynamic environments. The policy should adapt to two potentially conflicting navigation objectives: maximizing the visibility of a visual object as a perceptual constraint and obstacle avoidance to ensure safe flight.
-
Event Representation Learning for Control with Visual Distractors

This project develops event-based representation learning methods for control tasks in environments with visual distractors, leveraging sparse, high-temporal-resolution event data to improve robustness and efficiency over traditional frame-based approaches.
-
Rethinking RNNs for Neuromorphic Computing and Event-based Vision

This thesis develops hardware-optimized recurrent neural network architectures with novel parallelization and kernel-level strategies to efficiently process event-based vision data for real-time neuromorphic and GPU-based applications.
-
Time-continuous Facial Motion Capture Using Event Cameras

Traditional facial motion capture systems, including marker-based methods and multi-camera rigs, often struggle to capture fine details such as micro-expressions and subtle wrinkles. While learning-based techniques using monocular RGB images have improved tracking fidelity, their temporal resolution remains limited by conventional camera frame rates. Event-based cameras present a compelling solution, offering superior temporal resolution without the cost and complexity of high-speed RGB cameras. This project explores the potential of event-based cameras to enhance facial motion tracking, enabling the precise capture of subtle facial dynamics over time.
-
Vision-Based Tactile Sensor for Humanoid Hands (in collaboration with Soft Robotics Lab)

Humanoid robots require tactile sensing to achieve robust dexterous manipulation beyond the limits of vision-based perception. This project develops an event-based tactile sensor to provide low-power, high-bandwidth force estimation from material deformation, with the goal of integrating it into a human-scale robotic hand.
-
Vision Language Action models for Drones

This project explores generative modeling of drone flight paths conditioned on natural language commands and spatial constraints, aiming to produce plausible 3D trajectories for training reinforcement learning policies. We investigate model architectures, data sources, and trajectory extraction methods to ensure generated paths are both physically feasible and stylistically aligned with textual descriptions.
-
Self-Supervised Event-Driven World Models for High-Speed Scene Forecasting

Predicting how a dynamic scene will evolve is a cornerstone of safe, agile robotic navigation. While conventional self-supervised world models rely on frame-based video prediction, they fail during rapid motions due to motion blur and low sampling rates. Event cameras circumvent these limitations by tracking continuous brightness changes with microsecond latency. This project focuses on developing a Self-Supervised Event-Driven World Model (S-EWM) that learns to forecast future environmental states directly from raw, asynchronous event streams. By predicting future event distributions or synthesized frames without human labels, the model will capture the underlying physics of highly dynamic environments, serving as a powerful representation for downstream robotic perception.
-
Independent Moving Object Segmentation with the Aeveon Sensor

This project extends the Motion-aware Event Suppression framework by integrating multi-bit event streams and synchronized intensity data from the new iniVation Aeveon sensor to improve real-time object segmentation. It aims to develop a multi-modal neural architecture that disentangles ego-motion from independently moving objects, evaluated through real-world benchmark sequences.
-
Hardware-Aware Mapping and Quantization of Event-Driven GNNs for Edge Vision

This project focuses on the hardware-aware mapping and quantization of Asynchronous Graph Neural Networks (AGNNs) for edge vision applications. By optimizing DaGR network operations and developing advanced quantization techniques, the research aims to bridge the gap between software simulation and physical execution on the GNN processor, ensuring high performance under strict bit-width constraints.
-
Spiking Neural Networks for Agile Autonomous Drone Racing

Explore the use of Spiking Neural Networks (SNNs) for agile, autonomous drone racing, focusing on achieving low-latency and energy-efficient flight control. By leveraging event-based processing, the research aims to develop and validate high-performance spiking policies suitable for deployment on neuromorphic hardware.