Project ideas (UG/PGT)

Read this first

If you contact me and it is obvious from your e-mail that you did not read this page, I will ignore your request and I will not feel bad about it. If you have some idea in mind, read up on this short, less than perfect, but better than nothing guide to proposing your own project idea. If you do not, scroll down to the list of projects I have some interest in supervising. Some are better thought out than others, but I expect you to bring your ideas in them as well. I am not the kind of supervisor who gives weekly to-do lists.

List of potential projects

Here are a bunch of project ideas I would like to supervise in some form. They are not fixed in so far as you can come up with a slight variation of them and we can talk them out. They can be done at the undergraduate or the MSc level but might require slight adaptations in some cases to better fit the timeline of your degree. Some of those are research-oriented, and would fit well a student aiming for further study. Some are more engineering-focused, and would fit well a student who wants to build something cool (hopefully). I classify projects into three wide categories: (1) projects somewhat affiliated to my research group and might build on one another ; (2) one-shot projects which I think are fun and interesting but completely unrelated to my work ; (3) general lines of investigation that I am interested in, but without a clear direction (you will be expected to bring a lot more of your ideas into this).

The following projects are ongoing lines of interest for me. They might have been done to some extent by previous students, but that does not mean you can’t propose to give your own twist on it during our initial discussion phase.

Theme 1. Agents, Skills and Autonomous Systems

Project 1.1 Measuring the complexity and resource requirements of LLM skills

LLM-based agents increasingly make use of reusable skills or procedures that describe how particular tasks should be completed. However, different skills can have very different computational requirements: some may involve a short local procedure, while others require multiple model calls, external tools, significant context, or lengthy execution. In this project, you will investigate how the complexity and resource requirements of an LLM skill can be described and measured. You may consider factors such as execution time, token usage, number of reasoning or tool-use steps, cost, reliability, required hardware, and external dependencies. The project could potentially involve developing a profiling framework, predicting the cost of previously unseen skills, or studying which properties of a skill provide useful estimates of its practical complexity.

Tentative outcome: A software framework for profiling LLM skills, accompanied by an experimental research study evaluating different measures of skill complexity.

Project 1.2 Resource-aware skill selection for LLM agents

An intelligent agent may have several different ways of completing the same task, but the best choice can depend on the circumstances. A user may require a fast response, have a limited computational or monetary budget, or require that sensitive information remains on a local machine. In this project, you will build and evaluate an LLM-based agent that selects between alternative skills or procedures according to such constraints. The project will involve representing the capabilities and requirements of different skills and developing a mechanism that allows an agent to select an appropriate option for each task. Depending on your interests, the project could focus on latency, cost, privacy and execution boundaries, reliability, local versus remote models, or combinations of these constraints.

Tentative outcome: A working agent or routing system, supported by an empirical evaluation of its ability to satisfy different resource and privacy constraints.

Project 1.3 Progressive discovery of skills for LLM agents

As the number of skills available to an LLM-based agent increases, it becomes increasingly impractical to provide the complete description and instructions for every possible skill within the model’s context. In this project, you will investigate mechanisms that allow an agent to progressively discover the information it needs to select an appropriate skill. A system might initially expose only a compact description of each capability, retrieve additional information about promising candidates when required, and load the complete procedure only after a skill has been selected. You will develop and evaluate one or more approaches to this problem, examining the trade-off between selection accuracy and the amount of information the agent must process. Possible directions include information retrieval, hierarchical catalogues, structured skill metadata, adaptive disclosure, and large-scale skill selection.

Tentative outcome: A prototype skill-discovery system and a research evaluation comparing different mechanisms for representing and retrieving skills.

Project 1.4 An LLM committee that knows when not to agree

Multi-agent LLM systems often attempt to improve answers by allowing several agents to discuss a problem and converge on a common conclusion. However, consensus is not necessarily desirable when agents possess conflicting evidence or when an initially correct agent is persuaded by an incorrect majority. In this project, you will investigate alternative mechanisms for collective decision making in which disagreement can be preserved, quantified, or used to decide when a system should abstain or request additional information. You will build and evaluate a multi-agent system and study how different communication and decision protocols affect accuracy, computational cost, and the reliability of the final answer. The scale of the system can be adapted to the available hardware.

Tentative outcome: An extension to a multi-agent LLM platform, with a paper-style experimental comparison of alternative decision-making protocols.

Project 1.5 Protecting LLM agents from malicious or inappropriate instructions

LLM-based agents may receive information from documents, websites, databases, or other external sources. Some of this information may contain instructions that conflict with the user’s intentions or attempt to make the agent perform inappropriate actions. In this project, you will investigate mechanisms that allow an agent to distinguish between information it should use and instructions it is authorised to follow. You will build a small agent with access to a controlled set of tools and evaluate one or more approaches for enforcing permissions or execution boundaries. The project could focus on prompt injection, information-flow restrictions, tool permissions, local versus external execution, or the design of safer agent architectures.

Tentative outcome: A secure agent prototype or sandbox, together with a benchmark-based evaluation of one or more defensive mechanisms.

Project 1.6 Trading model complexity for harness complexity in LLM agents

Modern AI agents combine a language model with a software harness that manages tasks such as planning, tool use, state, validation, and error recovery. More capable models can perform many of these functions themselves, but they also require greater computational resources. In this project, you will investigate whether some of this capability can instead be moved into deterministic software, allowing smaller or more aggressively quantised models to achieve similar task performance. Depending on your interests and available hardware, the project could compare tiny language models, highly compressed models, and larger baselines across agentic tasks with different levels of harness support. The aim is to understand when additional system structure can compensate for reduced model capability, and what trade-offs this introduces in performance, latency, resource use, and generality.

Tentative outcome: A configurable agent harness and a research paper-style experimental evaluation of the trade-off between model capability and harness complexity.

Theme 2. Human-AI Interaction, Trust and Personalisation

Project 2.1 An agent that knows when not to use what it remembers

Long-term memory can make conversational agents more useful and personalised, but information that is appropriate in one context may be inappropriate in another. In this project, you will investigate mechanisms that allow users or agents to control not only what information is remembered, but also when that information may be retrieved and used. You will design a memory architecture and prototype conversational agent that supports contextual restrictions or permissions around stored information. Possible directions include user-defined memory boundaries, automatic detection of sensitive contexts, memory provenance, inspectable memory controls, and the effect of such mechanisms on usefulness, privacy, and user trust.

Tentative outcome: A conversational software prototype with contextual memory controls, potentially accompanied by a user study or controlled experimental evaluation.

Project 2.2 An LLM assistant that knows when to ask for clarification

Large language models frequently receive requests that are incomplete or ambiguous. An assistant can either make an assumption and continue, or interrupt the interaction to request additional information. Both behaviours have costs. In this project, you will investigate when an AI assistant should ask a clarification question and what information that question should request. You will build an interactive system that can decide whether sufficient information is available to complete a task and evaluate different clarification strategies. Depending on your interests, the project could focus on uncertainty estimation, dialogue management, information gain, user effort, task success, or applications within a particular domain.

Tentative outcome: An interactive conversational prototype and an experimental study comparing alternative clarification strategies.

Project 2.3 An AI assistant that can disagree appropriately

Conversational AI systems are often designed to appear helpful and agreeable, but excessive agreement can cause them to reinforce incorrect assumptions made by their users. In this project, you will investigate how an AI assistant can identify when it should challenge or correct a user while maintaining an appropriate conversational style. You will develop and evaluate one or more mechanisms for separating factual judgement from conversational behaviour. Possible applications include programming assistance, education, decision support, or other domains in which users may make incorrect assumptions.

Tentative outcome: A prototype conversational system and a research paper-style evaluation of factual correctness, disagreement behaviour, and potentially user perception.

Project 2.4 Repairing trust after an AI makes a mistake

Errors are unavoidable in deployed AI systems, but relatively little is understood about what an AI should do after an error has occurred. In this project, you will investigate different mechanisms through which an AI system can acknowledge, explain, and correct a previous mistake. You will build an experimental interface in which users interact with an AI assistant and compare different forms of error recovery, such as correction, explanation, evidence, or explicit acknowledgement of uncertainty. The exact application domain can be chosen according to the student’s interests.

Tentative outcome: Primarily a human-AI interaction research study, supported by an experimental software interface and suitable for presentation in research-paper form.

Project 2.5 An AI tutor that knows when not to provide the answer

Generative AI can provide students with immediate solutions to difficult problems, but immediate answers may not always produce effective learning. In this project, you will design an AI tutoring system that decides how much assistance to provide and when to provide it. Rather than always supplying a complete solution, the system might request an attempt, provide a hint, ask the student to explain their reasoning, or gradually reveal additional information. You will evaluate how different assistance strategies affect task completion, learning, or user experience. The project can be applied to programming, mathematics, data science, or another suitable subject area.

Tentative outcome: An AI tutoring prototype and a user study or learning experiment comparing different assistance policies.

Theme 3. Brain Data, Mental Workload and Adaptive Interfaces

Project 3.1 Learning useful representations from fNIRS data

Large fNIRS datasets create opportunities to investigate whether machine learning models can learn useful representations of brain activity without requiring large quantities of carefully labelled data. In this project, you will explore techniques for learning from fNIRS recordings and investigate whether the resulting representations can improve performance on downstream tasks such as mental workload classification. Depending on your interests, the project could focus on self-supervised learning, transfer learning, model comparison, robustness across participants or studies, or reducing the quantity of labelled data required to train an effective model. The precise modelling approach and datasets used will depend on your machine learning background and available computational resources.

Tentative outcome: A trained machine learning model and experimental research paper examining representation learning or transfer between fNIRS datasets.

Project 3.2 Assessing the quality and reliability of fNIRS signals

Physiological sensing systems need to determine whether the data they receive are sufficiently reliable before using them to make decisions about a user. In this project, you will investigate automated methods for assessing the quality of fNIRS signals. You will develop and evaluate one or more machine learning approaches that attempt to distinguish reliable from problematic signal segments, with particular attention to whether models generalise across participants, experiments, or datasets. Depending on your interests, the project could emphasise interpretable signal features, deep learning, uncertainty estimation, visualisation, or the development of a small signal-inspection application.

Tentative outcome: A machine learning model for signal-quality assessment, potentially accompanied by a small visualisation or inspection tool and an empirical evaluation.

Project 3.3 A mental-workload classifier that knows when it does not know

Machine learning systems that infer mental workload from physiological signals will sometimes encounter users or signals that differ substantially from their training data. In such cases, producing a confident prediction may be less useful than recognising uncertainty. In this project, you will investigate uncertainty-aware approaches to mental workload classification using fNIRS or related physiological data. You will build and compare classifiers that can estimate the reliability of their predictions and potentially abstain when confidence is insufficient. The project could explore calibration, participant-to-participant generalisation, out-of-distribution detection, selective prediction, or the design of interfaces that communicate uncertainty to users.

Tentative outcome: An uncertainty-aware classification model and a research paper-style evaluation of calibration, generalisation, and selective prediction.