What is the solubility of a molecule? How do certain genes interact with diseases? Which movies might users prefer based on their profiles? How do proteins fold into their native 3D structures? How can we accurately estimate arrival times on road networks?

These diverse questions share a common thread: they all require machine learning on structured, relational data, including graphs (conventional or geometric), knowledge bases (such as knowledge graphs and databases), and other relational representations. Such data is deeply embedded across domains, including the life sciences, and forms the backbone of many high-impact real-world systems.

We focus on advancing machine learning methods for relational data. Traditionally, this has involved developing and analysing models such as graph neural networks and graph transformers. More recently, our work has shifted towards foundation models for relational data: large-scale, pre-trained models that aim to replicate the success of large language models in the graph domain. Unlike task-specific methods, these models are designed to generalise across tasks and domains, making them more suitable for real-world scientific and industrial applications.

A central goal is to theoretically characterise the capabilities and limitations of existing methods, particularly in terms of expressiveness, generalisation, and transferability, and to use these insights to design novel architectures from first principles. Ultimately, we aim to apply these next-generation models to high-impact scientific challenges, making graph-based machine learning more interpretable, scalable, and reliable, especially in biology, chemistry, and physics.

Key Directions

  • Graph foundation models: Single models that transfer across graphs, feature spaces, and tasks, built on equivariance and on shared interfaces such as random walks.
  • Knowledge graphs: Inductive and zero-shot link prediction over knowledge graphs and relational hypergraphs, knowledge base completion with box embeddings, and complex query answering.
  • Relational deep learning: Learning from relational databases and tables, including LLM agents that write interpretable SQL feature programs, and tabular foundation models for graphs.
  • Expressive power and theory: Characterising graph neural networks through Weisfeiler-Leman tests, homomorphism counts, random node initialisation, and zero-one laws.
  • New architectures: Cooperative GNNs, shortest path networks, planar graph networks, and learning on large graphs via intersecting communities.
graph-foundation-modelsknowledge-graphsrelational-databasesgnn-theory

People

Selected Publications

All publications →