Workshop on Formal Verification for Stochastic Analysis and Quantitative Finance
A one-day workshop on using formal verification for mathematical research and software engineering, focusing on mathematical and computational methods in quantitative finance.
Background
Earlier this summer, Kevin Buzzard ran a workshop where a group of Lean users experimented with AI tools for writing and checking formal proofs, much of it related to Kevin’s project on Fermat’s Last Theorem. Logos sponsored the event and used the occasion to release a beta version of our harness. It was a great week, and it left us keen to explore how these tools might be used outside the formalisation community.
This is the sequel workshop, this time targeted at researchers and engineers in stochastic analysis and quantitative finance. No prior experience with Lean or formal proof assistants is required.
The plan
The plan is to investigate mathematical artefacts drawn from participants’ own work, the literature, or AI-generated. These might include numerical algorithms, optimisation procedures, simulation methods etc. The goal is to explore how Logos Harness can help in identifying mistakes, finding counterexamples, verifying correctness, checking assumptions, translating, or migrating code between languages and frameworks while preserving behaviour, or even suggesting alternative formulations and new algorithmic ideas. We hope to explore a broad spectrum of problems, ranging from stochastic and partial differential equations to convergence and stability of numerical schemes, efficient simulation algorithms, optimisation methods, legacy code migration, and algorithm discovery more broadly.
We’ll also dedicate time to a wider moderated discussion on the risks posed by AI in mathematics and finance and strategies for better controlling AI agents in accuracy-critical domains.
Participants
Alexander Barzykin
Director in FX, Rates and Commodities
HSBC
Christian Bayer
Principal Researcher in Stochastic Algorithms and Nonparametric Statistics
WIAS Berlin
Claudio Bellani
Quant Strategist
Citadel
Nick Chambers
Head of OTC Engineering
Citadel
Nicola Muca Cirone
Member of Technical Staff
Cartesia
Peter Friz
Einstein Professor in Mathematics
TU Berlin
Paolo Guasoni
Professor of Mathematical Finance
Dublin City University
Massimiliano Gubinelli
Wallis Professor of Mathematics
University of Oxford
Tom Hadfield
Senior Research Fellow in Quantitative Finance
Imperial College London
Martin Hairer
Professor of Pure Mathematics
EPFL and Imperial College London
Blanka Horvath
Professor of Mathematical and Computational Finance
University of Oxford
Antoine Jacquier
Professor of Mathematical Finance
Imperial College London
Tiger (Yuhao) L.
AI Engineer
HSBC
Gordon Lee
Head of Markets Quants
BNY
Johannes Muhle-Karbe
Head of Mathematical Finance
Imperial College London
Director of the CFM-Imperial Institute of Quantitative Finance
Eyal Neuman
Professor of Mathematical Finance
Imperial College London
Hao Ni
Professor of Mathematics
University College London
Roel Oomen
Head of Quant R&D
Deutsche Bank
Grigoris Pavliotis
Professor of Applied Mathematics
Imperial College London
Alexander Povey
Quantitative Researcher
Citadel
Andreas Sojmark
Professor of Mathematical Finance
London School of Economics
Lukasz Szpruch
Professor of Mathematics
University of Edinburgh
Director for Finance and Economics at The Alan Turing Institute
Edoardo Vittori
Business Director
Intesa Sanpaolo
Yufei Zhang
Professor of Mathematical Finance
Imperial College London
Programme
Timings are indicative and each talk includes time for questions.
- 09:50Introductory remarks
- 10:00Introduction to LeanJustus SpringerIn light of recent advances in AI for mathematics, I will discuss the role of Lean and formalisation in the current environment. I will discuss its benefits in terms of trustworthiness and increased mathematical understanding, and how Lean is used currently to certify proofs found using generative AI.
- 10:25Experiments in Verified ProgrammingDavid LedvinkaIn this talk I will discuss several experiments in verified programming and potential industry applications such as verified code migration and verified algorithms for numerical computation. We will discuss the findings of the experiments as well as potential for future work.
- 10:50Progress in Formalisation: From Stochastic Integration to Option PricingTom KloseIn the first part of the talk, I will report on ongoing public efforts to formalise stochastic integration in Lean, give a personal view on the difficulties, and show first evidence that the pathwise approach presents a promising alternative route for formalisation. The main focus will be on the second part of the talk, in which I will explain how the Logos platform can be used to detect the little Heston trap, a known problem in mathematical finance that led to the mispricing of European options.
- 11:40Coffee
- 12:00Verifying Numerical Software in LeanNikolas TapiaNumerical software sits between exact mathematical intent and finite-precision execution. This talk presents our approach to recovering specifications from existing code, proving properties in Lean 4, and testing specifications against source implementations. Using QuantLib as a case study, we show how proofs and differential fuzzing provide complementary evidence, where floating-point error bounds remain necessary, and what is needed to scale this process.
- 12:25Math in the era of language machinesMassimiliano GubinelliMy talk will have two sides: on one side I will report on some experiments I’ve performed with Logos platform in the formalisation of certain aspects of probability and analysis. On the other side I want to spell out some general comments which I think relevant to the discussion on the role of AI in Math.
- 13:05Lunch
- 14:05Generating formally verified GPU kernelsArchie BrowneThe problem of translating high level, synchronous code to lower level, parallel code is extremely important in the design of efficient GPU kernels. Developers would value the ability to automatically compile high-level neural network code, written in for example PyTorch, to more efficient CUDA code in a verified manner. Benchmarks such as KernelBench task LLMs with making that translation but provide no formal guarantees on the semantic equivalence of LLM-generated code with the PyTorch specification, relying on basic testing. Nor do they ensure that the generated code is memory or thread safe, which is highly important when designing parallel code. Recent additions to the Lean ecosystem such as Velvet have made it possible to reason formally about imperative code with pre/post conditions and invariants handed to an automatic proof discharger such as Z3 in a Dafny-like manner. This talk describes how we have made use of these libraries in order to provide an environment in Lean for proving the semantic equivalence between kernels written in PyTorch and CUDA. We tasked an LLM with generating some basic kernels and we will review some examples of where this environment caught semantic discrepancies. Time permitting, we will also review how it is possible to not only prove semantic equivalence, but also thread and memory safety of the generated code without executing it. Finally, we will conclude by showing how our experiments motivate the need for better floating point support in Lean.
- 14:30To be confirmedMartin Hairer
- 14:55OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant AgentsHao NiIn this talk, I will introduce OpenFinGym, a unified environment for developing and evaluating quantitative-finance agents across forecasting, market generation, real-time trading, and fraud detection. OpenFinGym combines a common execution and verification interface with an automated pipeline that converts research papers into executable tasks, a secure containerised runtime designed to prevent train–test leakage, and a low-latency paper-trading engine. It also supports deferred evaluation of long-horizon forecasts and integration with supervised fine-tuning and reinforcement learning for training quant agents.
- 15:20Coffee
- 15:40Moderated discussionAll participants