Interpretability for LLMs

This lecture focuses on the other way of approaching the problem—the speculative interpretability work that Anthropic and OpenAI have published on distill.pub. Despite its somewhat shaky theoretical foundations and aggressive duck typing, the results are striking. With a few caveats—we must never forget the speculative assumptions involved—these techniques point toward a new generation of networks that are lightweight and accurate. The Qwen models seem to be among the first heading in precisely that direction.

Raphael Korsoski

Developer, educator and independent researcher in computer science and machine learning