Thinking or Outsourcing? Comparing a Specialized GenAI Tool with a General-Purpose Chatbot
Motivation & Problem
General-purpose chatbots such as ChatGPT or Claude answer quickly. That is their appeal — and, for complex problem solving, potentially their problem. When plausible solutions arrive immediately, users may shift from working a problem through to merely reviewing and lightly editing what the system produced. Recent work describes this as shallow engagement and links it to cognitive offloading: the reasoning that makes the task valuable in the first place gets handed to the machine.
An alternative design ambition is productive engagement: the user draws on AI support but remains the driver of problem formulation and reasoning. Specialized GenAI tools are emerging that pursue exactly this — deliberately keeping users in the thinking rather than handing them answers. Whether they succeed, however, is largely untested: the relevant comparison is not against working unaided, but against the general-purpose chatbot users would otherwise reach for.
The answer may well be a trade-off rather than a win. A chatbot likely produces a polished output faster; a tool that keeps users thinking likely produces more thinking. Whether deeper engagement also yields a better outcome — or whether the two come apart — is what this thesis sets out to measure. The overarching question is: How does a specialized GenAI tool, compared to a general-purpose chatbot, affect users’ critical engagement during problem solving and the quality of the resulting work outcome?
The thesis does not start from scratch. It can build on prior work in our group, including a prior qualitative study and access to a specialized tool through a startup we collaborate with. Building on this foundation, you will develop a controlled comparison that examines both a process level (how users engage with the task) and an outcome level (the quality of what they produce), so that engagement and performance can be compared against each other rather than assumed to move together.
What You Will Do
- Review the literature on GenAI-supported problem solving, cognitive offloading, and critical thinking, and derive hypotheses for the comparison.
- Design the experiment. Develop an experimental design and select measures for the process level and the outcome level.
- Run the study with participants, following ethical and data-protection requirements (informed consent, debriefing).
- Analyze the data quantitatively — comparing conditions on process and outcome measures and examining how the two relate.
- Write up the findings, discussing what they mean for the design of GenAI tools that aim to support thinking rather than replace it.
Your Profile
- Genuine interest in empirical, experimental research on human–AI collaboration and how AI shapes thinking at work
- Willingness to work with data in R, Python, or SPSS
- Interest in conducting experiments and in working with an existing startup
- Nice to have (not required): prior exposure to experimental methods or survey design
Please get in touch with a short email including your CV, a current transcript of records, a few sentences on why this topic interests you, and the planned start/finish date: hise@ifi.uzh.ch