Bachelor/Master Theses and Master Project Topics
This pages lists the open BSc. and MSc. thesis descriptions, as well as the master projects opportunities currently available in the DDIS research group.
If you are interested in any of the listed projects, please do not hesitate to contact the person mentioned in the open topic description.
If there are currently no open topics but you are generally interested in our research (see https://www.ifi.uzh.ch/en/ddis/research.html), or if you would like to propose a thesis about your own idea, you can send us an email to ddis-theses@ifi.uzh.ch.
Master Thesis: Generating Realistic Contact Networks for Epidemic Modelling
The COVID-19 pandemic showed how quickly an infectious disease can spread, and many experts consider further pandemics likely. To be better prepared next time, researchers and public health authorities need models that capture how a disease spreads from person to person. Such models require data on who was in contact with whom, which is among the most sensitive data there is. The Swiss COVID app, based on a privacy-preserving protocol, was deliberately designed so that authorities could not collect individual contact data through the app. This protects privacy, but it also leaves researchers with hardly any realistic contact data to develop and test their methods on. The few empirical contact networks that are publicly available are usually small, with a few hundred people, and cover only a narrow part of daily life, such as a single office or a group of students.
The goal of this thesis is to review and develop methods to generate realistic (synthetic) contact networks that can be used for epidemic modelling. This includes implementing existing approaches and comparing them systematically. Aspects that are important when synthetically generating realistic contact networks include: (1) temporality, i.e., knowing not only that two people were in contact, but also when; (2) scale, i.e., whether a generator can produce networks of realistic size for an epidemic, with hundreds of thousands or millions of people; and (3) structural properties such as burstiness (many contacts within a short time, e.g., during commuting hours), hubs (e.g., train stations) and stratification (e.g., younger people tend to mix mostly among themselves). These are examples only; the thesis is not expected to cover all of them.
Requirements
- Strong Python skills
- Basics in Statistics
Recommended Courses
- Network Science
- Mathematics for AI and Data Science
Thesis period: 07.01.2027 - 07.07.2027 (flexible)
Application deadline: 30.11.2026 (flexible)
The thesis is embedded in our research group at the University of Applied Sciences and Arts Northwestern Switzerland (FHNW) in Olten, and the student can collaborate with its members, including Dr. Martin Sterchi and Dr. Lorenz Hilfiker. Occasional attendance at meetings in Olten is welcome but not required.
Throughout the thesis, the student will be assisted by me, Nathan Brack, a PhD student in the group. For more information or to apply, please contact me at nathan.brack@uzh.ch. Please note that I am on holiday from 10 to 18 October 2026 and will not answer emails during that time.
Master Thesis: Active Learning for Rare Disease Diagnosis
A disease is considered rare when it affects fewer than 1 in 2,000 people. More than 6,000 such conditions are known, and although each one is rare on its own, together they affect an estimated 300 million people worldwide. Because any individual clinician encounters very few cases of a given condition, patients in Europe wait close to five years on average before receiving a correct diagnosis. Reaching a diagnosis depends on the patient's phenotypes, i.e., the observable signs and symptoms of a condition, such as seizures, short stature, or muscle weakness.
The goal of this thesis is to make a rare disease retrieval model correct itself from targeted feedback. Systems that match a patient's phenotype profile against candidate diseases typically produce a single static ranking that cannot be corrected once it has been produced. The thesis builds on factorized models from recent work that learn, for each rare disease, a vector indicating which similarity metrics to trust. The aim of the thesis is to design and implement an active learning loop around it: the candidate diseases the model is least certain about, for example, those separated by the smallest score margin, are selected for feedback, and that feedback updates the per-disease weights. In a second step, the annotator is simulated with a Large Language Model (LLM), and retrieval performance (Hit@k, MRR, nDCG) is tracked across feedback rounds against a random-sampling strategy and a static baseline, together with an analysis of when LLM feedback helps and when it introduces errors. Feedback from human experts is planned as future work and is out of scope for this thesis.
Requirements
- Strong Python skills
- Solid background in machine learning
- Experience running and prompting Large Language Models
- Knowledge of ontologies or knowledge graphs (bonus)
- No medical background required
Recommended Courses
- Advanced Topics in Artificial Intelligence
- Advanced Machine Learning
- Deep Learning
Thesis period: Spring Semester 2027
Application deadline: 30.11.2026
For more information about the thesis, contact Pascal Andermattpandermatt@ifi.uzh.ch
Master Thesis: Graph-Based Patient Similarity for Rare Disease Diagnosis
A disease is considered rare when it affects fewer than 1 in 2,000 people. More than 6,000 such conditions are known, and although each one is rare on its own, together they affect an estimated 300 million people worldwide. Reaching a diagnosis depends on the patient's phenotypes, that is the observable signs and symptoms of a condition, such as seizures or muscle weakness. Patient's phenotypes are compared to other patients (patient-to-patient similarities) or to disease definitions (patient-to-disease similarities). Much of the patient's phenotype information, however, is buried in unstructured clinical notes.
The goal of this thesis is to compare similarity methods that are based on flat lists of phenotypes to methods that consider structural features such as patient-centric graphs. For this, each patient's medical history needs to be represented as a knowledge graph built from their clinical notes. Nodes in the graph represent phenotypes, extracted from clinical notes with existing named entity recognition (NER) methods, for example, domain-trained neural classifiers refined with Large Language Model (LLM) filtering. Edges in the graph represent contextual information about the phenotypes, such as whom the phenotype belongs to, whether it is negated, and when it occurred.
Given the patient-centric structured representations, graph embeddings will be generated per patient using graph neural networks (GNNs). The learning algorithms to be tested include, but are not limited to, Relational Graph Attention Networks (rGAT), Relational Graph Convolutional Networks (rGCN), and Heterogeneous Graph Transformer (HGT). The approach will be evaluated in two stages: first patient-to-patient similarity on the full MIMIC-IV cohort, then patient-to-patient and patient-to-disease similarity restricted to the smaller subset of MIMIC-IV patients with a documented rare disease.
Requirements
- Strong Python skills
- Solid background in natural language processing and machine learning, in particular graph representation learning (GNNs, knowledge graph embeddings)
- Familiarity with knowledge graphs and ontologies
- Experience running and prompting Large Language Models
- No medical background required
Recommended Courses
- Advanced Topics in Artificial Intelligence
- Advanced Machine Learning
- Machine Learning for Natural Language Processing 1
Thesis period: Spring Semester 2027
Application deadline: 30.11.2026
For more information about the thesis, contact Pascal Andermatt pandermatt@ifi.uzh.ch and Selene Baez Santamaria sbaez@ifi.uzh.ch
Master Thesis: An LLM Moderator Against Topic Drift in Multi-Agent Political Debates
Political debate shows on television, such as SRF Arena, follow a familiar format: representatives of different parties argue over a single question, while a moderator keeps the discussion on topic, cuts off digressions and introduces new angles when the debate stalls. In our existing Debate-Bots project, we fine-tuned Large Language Models (LLMs) to hold specific political viewpoints and let them debate political topics in a format modelled on SRF Arena. The political alignment of the bots worked well, but as debates progress, the bots gradually drift away from the original topic. In human debates, this drift is usually prevented by the moderator. The artificial debates currently have no one in that role.
The goal of this thesis is to design, implement and evaluate a moderator agent for these multi-bot debates. The moderator follows the debate as it unfolds and decides when and how to intervene: for example, by interjecting when a contribution strays from the topic, by delivering cutting statements when a bot repeats itself or monopolises the discussion, or by posing new questions that steer the debate back to the original issue. Design choices include how the moderator detects drift (e.g., semantic similarity between contributions and the debate question over time), how often it should intervene, and how strongly its prompts constrain the debaters. The moderator is evaluated against an unmoderated baseline and a simple rule-based moderator that intervenes at fixed intervals. Evaluation covers three dimensions: (1) topic drift, measured over the course of a debate; (2) political alignment, i.e., whether the debaters keep their assigned viewpoints despite moderator interventions; and (3) realism and engagement of the debates, assessed for example with LLM-based judges or a small human rating study.
Requirements
- Strong Python skills
- Experience running, prompting and finetuning Large Language Models
Nice to Have
- Knowledge of Swiss German or German (as debates are based on SRF material)
- Interest in politics and political communication
Thesis period: Spring Semester 2027
Application deadline: 30.10.2026
For more information about the thesis, feel free to contact me at Daan van der Weijden
Master Thesis: Predicting Moderator Interventions from Deliberative Quality Over Time
Online deliberation platforms let citizens discuss political questions at a scale that was not possible before, but that same scale is outgrowing what human moderators can follow in real time. Existing moderation tools mostly detect toxicity, such as insults or hate speech. Many moderator interventions, however, respond to something subtler: a gradual decline in the quality of the discussion, for example when participants stop giving reasons, ignore each other's arguments, or lose respect for opposing views. By the time a discussion becomes toxic, the moment for a constructive intervention has often passed.
The goal of this thesis is to test whether this decline in quality is measurable before a moderator intervenes, and whether it can be used to predict upcoming interventions. The thesis uses the WHoW dataset, which contains around 20,000 annotated moderation sentences. In a first step, the student conducts an exploratory analysis of established deliberative quality measures, such as the Discourse Quality Index (DQI) and AQUA, in the turns leading up to moderator interventions, and compares them to stretches of discussion without intervention. In a second step, the student builds a logistic regression classifier over sliding-window features that capture shifts in quality, and predicts whether an intervention is imminent. The classifier is evaluated against simple baselines, such as a toxicity-only model, with particular attention to which quality dimensions carry the predictive signal. The aim is a simple, interpretable early-warning signal for moderators rather than a black-box toxicity model.
Requirements
- Strong Python skills
- Basics in Statistics
- Basic knowledge of machine learning and natural language processing
Nice to Have
- Interest in political discussion and deliberation
Thesis period: Spring Semester 2027
Application deadline: 30.10.2026
For more information about the thesis, feel free to contact me at Daan van der Weijden