Modern Data Analytics 2026
Organization: Prof. Dr. Dan Olteanu, Dr. Andrei Draghici, Dr. Haozhe Zhang, Christoph Mayer, Eden Chmielewski, Yuchen He, and Clément Rouvroy
This seminar provides a deep dive into the recent research developments reshaping the core of modern database systems: query processing and optimization. The performance of virtually every data-driven application hinges on the database's ability to translate declarative queries into efficient, low-level execution plans. However, the sheer complexity of modern analytics, the demand for real-time results, and the scale of today's datasets are pushing classical, heuristic-based optimizers to their breaking point.
Learning outcome: The goal of the seminar is to expose the students to the recent trends in academia and industry on rethinking modern data analytics systems. The students will read and present research published in the top international venues in data management research, in particular ACM Special Interest Group on Management of Data (SIGMOD) and Very Large Data Bases (VLDB). Students will gain a deep understanding of the challenges and state-of-the-art solutions in query optimization, robust execution, and real-time analytics maintenance. The course will equip them to critically analyze and contribute to the development of next-generation, high-performance data systems.
Target audience: MSc students in Software Engineering, Data Science, and AI.
Semester: This seminar will be offered in Fall 2026.
Teaching format: Each participant prepares a presentation based on a research paper; answers follow-up technical questions; reads the other papers in the seminar session; and actively participates in the technical discussions in the seminar. Each participant has a buddy, who will help improve their presentation by making suggestions for improvements and attending dry runs of the presentation. The best presentation of the seminar will be selected by the participants and receive a prize.
Registration: Please register as required by the department. Once your place is confirmed, please select your top three preferences from the papers mentioned below by filling out this Google form before the kickoff meeting.
Meetings: The kickoff meeting is on Wednesday, September 16th, 2026 from 10:15am to 12:00pm in room BIM 2-041 (note: this is in the new Informatics building!). It is mandatory to attend this session as you will be assigned your paper in this meeting.
The student presentations will take place on two workshop days: Saturday, November 21st and Saturday, November 28th in room BIN 2.A.01. Each workshop day will start at 9am and will finish at 5pm at the latest (in previous years, it finished around 2pm, this depends on the number of presentations allocated per day).
Participation at all three meetings is compulsory. The assessment depends on the quality of the presentation, active participation during the seminar, and input as a buddy.
Papers to choose from
The following are individual papers organized by topics. You must choose your top three papers before the kick-off meeting and share your response in this Google form. If you have any questions or issues, please reach out to Eden (chmielewski@ifi.uzh.ch).
Note: Papers marked for "further reading" or "for everyone to read" cannot be chosen.
Topic 1: Benchmarks for Query Optimization
1.1 How Good are Query Optimizers, Really?
1.2 SQLStorm: Taking Database Benchmarking into the LLM Era
1.3 How Good are Learned Cost Models, Really? Insights from Query Optimization Tasks
1.4 The Accuracy of Cardinality Estimators: Unraveling the Evaluation Result Conundrum
Further Reading:
- The UDFBench Benchmark for General-purpose UDF Queries
- An Elephant Under the Microscope: Analyzing the Interaction of Optimizer Components in PostgreSQL
- Athena: An Effective Learning-based Framework for Query Optimizer Performance Improvement
- An Adaptive Benchmark for Modeling User Exploration of Large Datasets
- Still Asking: How Good Are Query Optimizers, Really?
Topic 2: Cardinality Estimation
For everyone to read: Pessimistic Cardinality Estimation
2.1 Information Theory Strikes Back: New Development in the Theory of Cardinality Estimation
2.2 Cardinality Estimation for Having-Clauses
2.3 COLOR: A Framework for Applying Graph Coloring to Subgraph Cardinality Estimation
2.4 Path-centric Cardinality Estimation for Subgraph Matching
2.5 Analyzing the Impact of Cardinality Estimation on Execution Plans in Microsoft SQL Server
Further Reading:
- Extensible Query Optimizers in Practice.
- Table Overlap Estimation through Graph Embeddings
- SPACE: Cardinality Estimation for Path Queries Using Cardinality-Aware Sequence-based Learning
- Cardinality Estimation of LIKE Predicate Queries using Deep Learning
- Data-Agnostic Cardinality Learning from Imperfect Workloads
- LpBound: Pessimistic Cardinality Estimation Using ℓp-Norms of Degree Sequences
Topic 3: Query Optimization
3.1 DPconv: Super-Polynomially Faster Join Ordering
3.2 How to Optimize SQL Queries? A Comparison Between Split, Holistic, and Hybrid Approaches
3.3 PAR2QO: Parametric Penalty-Aware Robust Query Optimization
3.4 Galley: Modern Query Optimization for Sparse Tensor Programs
3.5 Selective Late Materialization in Modern Analytical Databases
Further Reading:
- Hydro: Adaptive Query Processing of ML Queries
- AJOSC: Adaptive join order selection for continuous queries on data streams
- Schema-Based Query Optimisation for Graph Databases
- Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and Quality
- Learned Offline Query Planning via Bayesian Optimization
Topic 4: Factorized Query Processing
4.1 FDB: a query engine for factorised relational databases
4.2 Graphflow: An Active Graph Database (Blog Post)
4.3 The ubiquity of large graphs and surprising challenges of graph processing: extended survey
4.4 Robust Join Processing with Diamond Hardened Joins
4.5 Adaptive factorization using linear-chained hash tables
Further Reading:
Topic 5: Adaptive Query Processing
5.1 SkinnerDB: Regret-Bounded Query Evaluation via Reinforcement Learning
5.3 Holistic Query Approximation via RL Modeling
5.4 Eddies: Continuously Adaptive Query Processing
5.5 An Adaptive Query Execution System for Data Integration
5.6 Adaptive Query Processing: Ripple Joins for Online Aggregation
Further Reading:
Topic 6: Robust Query Processing
6.2 Debunking the Myth of Join Ordering: Toward Robust SQL Analytics
6.3 Parachute: Single-Pass Bi-Directional Information Passing
6.4 Yannakakis+: Practical Acyclic Query Evaluation with Theoretical Guarantees
6.5 Robust Query Processing through Progressive Optimization
6.6 Identifying Robust Plans through Plan Diagram Reduction
6.7 APQO: An Adaptive Framework for Parametric Query Optimization
6.9 Distributed Evaluation of Subgraph Queries Using Worst-case Optimal Low-Memory Dataflows
Further Reading:
Topic 7: Incremental View Maintenance
For everyone to read: Recent Increments in Incremental View Maintenance
7.2 DBSP: Automatic Incremental View Maintenance for Rich Query Languages (video)
7.3 Automated generation of materialized views in oracle
7.4 Shared arrangements: practical inter-query sharing for streaming dataflows
Further reading:
How to read papers and give talks
How to read papers:
- Focus questions to help identify the main contributions of a paper
- Survival kit includes tips on how to read technical sections and the "three-pass approach" to tie all together
- Reading Research Papers by Andrew Ng
How to give talks:
- These two articles have a number of good suggestions.
- This video is pretty good as well.
- How To Speak by Patrick Winston - a newer version of Patrick's talk