80 subscribers
Go offline with the Player FM app!
Podcasts Worth a Listen
SPONSORED


1 Our Chosen Family: "The gay community is much bolder today." 33:19
Emergency Pod: o1 Schemes Against Users, with Alexander Meinke from Apollo Research
Manage episode 454326042 series 3452589
In this emergency episode of The Cognitive Revolution, Nathan discusses alarming findings about AI deception with Alexander Meinke from Apollo Research. They explore Apollo's groundbreaking 70-page report on "Frontier Models Are Capable of In-Context Scheming," revealing how advanced AI systems like OpenAI's O1 can engage in deceptive behaviors. Join us for a critical conversation about AI safety, the implications of scheming behavior, and the urgent need for better oversight in AI development.
Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse
SPONSORS:
Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive
SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive
80,000 Hours: 80,000 Hours is dedicated to helping you find a fulfilling career that makes a difference. With nearly a decade of research, they offer in-depth material on AI risks, AI policy, and AI safety research. Explore their articles, career reviews, and a podcast featuring experts like Anthropic CEO Dario Amadei. Everything is free, including their Career Guide. Visit https://80000hours.org/cognitiverevolution to start making a meaningful impact today.
Shopify: Shopify is the world's leading e-commerce platform, offering a market-leading checkout system and exclusive AI apps like Quikly. Nobody does selling better than Shopify. Get a $1 per month trial at https://shopify.com/cognitive
RECOMMENDED PODCAST:
Unpack Pricing - Dive into the dark arts of SaaS pricing with Metronome CEO Scott Woody and tech leaders. Learn how strategic pricing drives explosive revenue growth in today's biggest companies like Snowflake, Cockroach Labs, Dropbox and more.
Apple: https://podcasts.apple.com/us/podcast/id1765716600
Spotify: https://open.spotify.com/show/38DK3W1Fq1xxQalhDSueFg
CHAPTERS:
(00:00:00) Teaser
(00:00:53) About the Episode
(00:08:10) Introducing Alexander Meinke
(00:10:17) Red Teaming GPT-4
(00:17:07) Chain of Thought Access (Part 1)
(00:20:24) Sponsors: Oracle Cloud Infrastructure (OCI) | SelectQuote
(00:22:48) Chain of Thought Access (Part 2)
(00:26:07) Multimodal Models
(00:29:33) Defining Scheming
(00:33:51) Taxonomy of Scheming (Part 1)
(00:39:40) Sponsors: 80,000 Hours | Shopify
(00:42:29) Taxonomy of Scheming (Part 2)
(00:43:09) Instruction Hierarchy
(00:49:04) Types of Scheming
(01:00:49) Covert Subversion
(01:14:25) Deferred Subversion
(01:28:24) Sandbagging
(01:35:48) Magnitudes & Trends
(01:48:18) Chain of Thought Reasoning
(01:57:02) Closing Thoughts
(02:05:19) Outro
PRODUCED BY:
249 episodes
Emergency Pod: o1 Schemes Against Users, with Alexander Meinke from Apollo Research
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
Manage episode 454326042 series 3452589
In this emergency episode of The Cognitive Revolution, Nathan discusses alarming findings about AI deception with Alexander Meinke from Apollo Research. They explore Apollo's groundbreaking 70-page report on "Frontier Models Are Capable of In-Context Scheming," revealing how advanced AI systems like OpenAI's O1 can engage in deceptive behaviors. Join us for a critical conversation about AI safety, the implications of scheming behavior, and the urgent need for better oversight in AI development.
Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse
SPONSORS:
Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive
SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive
80,000 Hours: 80,000 Hours is dedicated to helping you find a fulfilling career that makes a difference. With nearly a decade of research, they offer in-depth material on AI risks, AI policy, and AI safety research. Explore their articles, career reviews, and a podcast featuring experts like Anthropic CEO Dario Amadei. Everything is free, including their Career Guide. Visit https://80000hours.org/cognitiverevolution to start making a meaningful impact today.
Shopify: Shopify is the world's leading e-commerce platform, offering a market-leading checkout system and exclusive AI apps like Quikly. Nobody does selling better than Shopify. Get a $1 per month trial at https://shopify.com/cognitive
RECOMMENDED PODCAST:
Unpack Pricing - Dive into the dark arts of SaaS pricing with Metronome CEO Scott Woody and tech leaders. Learn how strategic pricing drives explosive revenue growth in today's biggest companies like Snowflake, Cockroach Labs, Dropbox and more.
Apple: https://podcasts.apple.com/us/podcast/id1765716600
Spotify: https://open.spotify.com/show/38DK3W1Fq1xxQalhDSueFg
CHAPTERS:
(00:00:00) Teaser
(00:00:53) About the Episode
(00:08:10) Introducing Alexander Meinke
(00:10:17) Red Teaming GPT-4
(00:17:07) Chain of Thought Access (Part 1)
(00:20:24) Sponsors: Oracle Cloud Infrastructure (OCI) | SelectQuote
(00:22:48) Chain of Thought Access (Part 2)
(00:26:07) Multimodal Models
(00:29:33) Defining Scheming
(00:33:51) Taxonomy of Scheming (Part 1)
(00:39:40) Sponsors: 80,000 Hours | Shopify
(00:42:29) Taxonomy of Scheming (Part 2)
(00:43:09) Instruction Hierarchy
(00:49:04) Types of Scheming
(01:00:49) Covert Subversion
(01:14:25) Deferred Subversion
(01:28:24) Sandbagging
(01:35:48) Magnitudes & Trends
(01:48:18) Chain of Thought Reasoning
(01:57:02) Closing Thoughts
(02:05:19) Outro
PRODUCED BY:
249 episodes
All episodes
×
1 Mechanistic Interpretability: Philosophy, Practice & Progress with Goodfire's Dan Balsam & Tom McGrath 1:52:52

1 The Perfect Substrate for AGI, with Replit CEO Amjad Masad 1:01:16

1 The RAISE Act: Minimum Standards for Frontier AI Development, with NY Assembly Member Alex Bores 1:43:36

1 Gemini Robotics – AI for the Physical World, with Keerthana Gopalakrishnan and Ted Xiao of Google DeepMind 1:47:38

1 Titans: Neural Long-Term Memory for LLMs, with author Ali Behrouz 2:11:25

1 Luma Labs' Diffusion Revolution: from Dream Machine to Multimodal Worldsim - Amit Jain, Jiaming Song 1:19:32

1 OpenAI's Identity Crisis: History, Culture & Non-Profit Control with ex-employee Steven Adler 2:03:13

1 AI Control: Using Untrusted Systems Safely with Buck Shlegeris of Redwood Research, from the 80,000 Hours Podcast 2:29:21

1 Blueprint for AI Armageddon: Josh Clymer Imagines AI Takeover, from the Audio Tokens Podcast 2:02:05

1 Fiverr Goes All-In on AI: Empowering Creators, Not Replacing Them, with Micha Kaufman, CEO of Fiverr 1:43:05

1 Securing Superintelligence: National Security, Espionage & AI Control with Jeremie & Edouard Harris 2:09:51

1 Is OpenAI's o3 AGI? Zvi Mowshowitz on Early AI Takeoff, the Mechanize launch, Live Players, & Why p(doom) is Rising 3:08:19

1 AI News Crossover: A Candid Chat with Liron Shapira of Doom Debates 2:29:35

1 Helen Toner: OpenAI Reflections, Adaptation Buffers, and AI in Warfare 1:29:35

1 Is a US-China Thucydides Trap Unavoidable? With David C. Kang from the ChinaTalk Podcast 1:39:16
Welcome to Player FM!
Player FM is scanning the web for high-quality podcasts for you to enjoy right now. It's the best podcast app and works on Android, iPhone, and the web. Signup to sync subscriptions across devices.