đź“‹Â Table of Contents
🔄 Reposted from Discovery Engines Substack, Dec 05, 2025

Who can I trust?
Imagine you’re eight years old, born into a wealthy family that owns a sprawling, highly lucrative enterprise. Suddenly you’re orphaned, and responsibility for this vast inheritance falls to you.
As a first step you - smartly - search for an adult to manage the estate until you’re older. You interview candidate after candidate but you keep running into the same question: how do you know who to trust? You want someone who genuinely prioritizes your long-term interests. Several appear sincere, yet it’s impossible to tell who’s truly aligned with you and who’s simply skilled at acting aligned. Or worse: someone who will obey your requests blindly even when those requests drive the company off a cliff.
It’s this analogy at the heart of “Why AI Alignment Could Be Hard with Modern Deep Learning”, a thought-provoking essay and that my learning group read as part of an AI Safety Fundamentals course I recently completed with MIT AI Alignment (MAIA).
In the spirited discussions that followed, we covered a number of strategies, from developing ways to “control” superior intelligences, to inspecting their thought process through techniques such as mechanistic interpretability and probing. My personal take-away: this stuff is really hard. There’s something inherently slippery about trying to control a mind that could be vastly more intelligent than our own, especially if the time-scales for control extend beyond our lifespans.
It’s these questions, and the larger landscape of AI safety, that set the theme for my recent sit-down with Riya Tyagi and Gatlen Culp, MIT undergrads and MAIA board members. In our conversation, we covered:
Fun fact: Riya is a Mechanistic Interpretability Researcher at the lab of Max Tegmark, well-known AI researcher, safety advocate, and author of Life 3.0, a book that first opened my eyes to the possibilities - and dangers - of superintelligent AI systems.
If you’re curious about the cutting edge of AI safety research, from the perspective of the next generation of AI researchers, this conversation is for you.