Understanding Interpretability
Let's dive into the details surrounding Interpretability. What's happening inside an AI model as it thinks? Why are AI models sycophantic, and why do they hallucinate? Are AI models ...
Key Takeaways about Interpretability
- AI models are trained and not directly programmed, so we don't understand how they do most of the things they do. Our new ...
- How can we use the language of causality to understand and edit the internal mechanisms of AI models? Atticus Geiger ...
- Neel Nanda from DeepMind presenting 'Mechanistic
- MIT 6.S897 Machine Learning for Healthcare, Spring 2019 Instructor: Peter Szolovits View the complete course: ...
- Adam Shai presented “Building the Science of
Detailed Analysis of Interpretability
A surprising fact about modern large language models is that nobody really knows how they work internally. At Anthropic, the ... Take your personal data back with Incogni! Use code WELCHLABS at the link below and get 60% off an annual plan: ... Lex Fridman Podcast full episode: https://www.youtube.com/watch?v=ugvHCXCOmm4 Thank you for listening ❤ Check out our ...
Art by @hamishdoodles Clipped from episode 19 of AXRP: https://youtu.be/3YbE7zybc5k?t=64 Transcript of that episode: ...
That wraps up our extensive overview of Interpretability.