The Most Forbidden Technique Zvi · lesswrong.com · published 2025-03-12 read 2026-08-05 length4,934 words tagsai-safety, interpretability, llms, reward-hacking, chain-of-thought links original · archive.org Logged without notes.