AI safety spans techniques for peering inside the "brains" of LLMs, like mechanistic interpretability, to platforms for pressure-testing agentic workflows before they reach production.

I see the case for it on two levels. Today, safety is what unlocks AI adoption in high-stakes but potentially transformational use-cases like healthcare, where a hallucination or data leak carries real consequences. And if the ASI prognosticators are right, the same work helps guard against existential risks.

Below are my explorations into the current state of the field, plus organizations I've interacted with that are doing important work.

📋 Table of Contents

🕯️ Deep-Dives

Building Safe AI – A conversation with MIT AI Alignment's Riya Tyagi and Gatlen Culp

🔄 Reposted from Discovery Engines Substack, Dec 05, 2025

Who can I trust?

Who can I trust?

Imagine you’re eight years old, born into a wealthy family that owns a sprawling, highly lucrative enterprise. Suddenly you’re orphaned, and responsibility for this vast inheritance falls to you.

As a first step you - smartly - search for an adult to manage the estate until you’re older. You interview candidate after candidate but you keep running into the same question: how do you know who to trust? You want someone who genuinely prioritizes your long-term interests. Several appear sincere, yet it’s impossible to tell who’s truly aligned with you and who’s simply skilled at acting aligned. Or worse: someone who will obey your requests blindly even when those requests drive the company off a cliff.

It’s this analogy at the heart of “Why AI Alignment Could Be Hard with Modern Deep Learning”, a thought-provoking essay and that my learning group read as part of an AI Safety Fundamentals course I recently completed with MIT AI Alignment (MAIA).

In the spirited discussions that followed, we covered a number of strategies, from developing ways to “control” superior intelligences, to inspecting their thought process through techniques such as mechanistic interpretability and probing. My personal take-away: this stuff is really hard. There’s something inherently slippery about trying to control a mind that could be vastly more intelligent than our own, especially if the time-scales for control extend beyond our lifespans.

It’s these questions, and the larger landscape of AI safety, that set the theme for my recent sit-down with Riya Tyagi and Gatlen Culp, MIT undergrads and MAIA board members. In our conversation, we covered:

Fun fact: Riya is a Mechanistic Interpretability Researcher at the lab of Max Tegmark, well-known AI researcher, safety advocate, and author of Life 3.0, a book that first opened my eyes to the possibilities - and dangers - of superintelligent AI systems.