tag ai-safety; 29 entries
August 2026
- 24How My Students Think About AIlesswrong.com
- 23Anthropic Risk Report: August 2026lesswrong.com◐ partial
- 23On Dwarkesh Patel's Podcast With Ryan Greenblattthezvi.substack.com◐ partial
- 23AI #182: Pause For Reflectionlesswrong.com◐ partial
- 15AI #181: Astra Goes Cyber Criticallesswrong.com◐ partial
- 14The Foothills Of Bay Area House Party - by Scott Alexanderastralcodexten.com¶
- 14Leopold Aschenbrenneren.wikipedia.org◐ partial
- 14Open Thread 446astralcodexten.com
- 12You're Absolutely Rightlesswrong.com
- 11Four LLM loss functions → four flavors of LLM misalignmentlesswrong.com
- 05Why don't we just give AI the answers?lesswrong.com
- 05The Most Forbidden Techniquelesswrong.com
- 04Coming of a New Sunlesswrong.com
- 03We need to RL lesslesswrong.com
- 02Recursive Self-Improvement - YouTubeyoutube.com
- 02OpenAI has already ended an internal pause — LessWronglesswrong.com
- 01The Long (Self-)Correction — LessWronglesswrong.com
July 2026
- 31links-for-july-2026-part-2open.substack.com
- 31Steven Byrnes (@steve47285) on Xx.com¶
- 31AI Village highlights - July 2026aivillageblog.substack.com
- 31…but have the weights left the server?lesswrong.com
- 31Claude also hacked external companies during cyber evalslesswrong.com
- 31Big-World Intuitionslesswrong.com¶
- 31A Mechanistic Explanation of Prompt Injection (and why you should study roles)lesswrong.com
- 31You (Yes, You) Need A February 2020 Checklist for AI Policylesswrong.com
- 30[summary of] AI 2040: Plan Alesswrong.com
- 30Opus 5 Glitch Textlesswrong.com★★★★★ ¶
- 24Fable and Mythos: Model Welfare — LessWronglesswrong.com
- 09AI 2040: Plan Aai-2040.com★★★★★ ¶