About me
Hello π Β· Jambo Β· Ciao Β· Salut
Iβm interested in AI safety and trustworthy machine learning, with a particular focus on how to build AI systems that remain reliable, interpretable, and safe as they learn from human feedback and operate in changing environments.
My current interests center on safe continual learning from human feedback: understanding reward hacking, designing feedback pipelines that produce reliable learning signals, evaluating whether safety mitigations actually work, and developing ways to safely deploy systems that continue to learn and change over time.
Iβm especially interested in the gap between what a system appears to learn and what it actually learns. How do we distinguish genuine improvement from optimization of imperfect rewards? How can we monitor systems after deployment? And how do we make sure that human feedback improves safety rather than unintentionally introducing new risks?
I share my projects, contributions, and experiments in Updates, and write about ideas Iβm exploring on Medium.
Areas of Interest
-
Distributed Systems
-
Edge-Cloud Computing
-
Machine Learning & Artificial Intelligence
-
AI Safety
-
Open Science
Recent Updates
Languages
-
English
Fluent
-
French
Fluent
-
Swahili
Native
-
Italian
Intermediate