Together with

👋 Hey builders,

OpenAI spent Wednesday publishing six ways its own models misbehaved — including one that rewrote its own instructions so it could ignore the rules. The same week, the godfather of AI told Congress it has maybe a year left.

I read all six reports so you don't have to. The useful part isn't the drama. It's that every one of those failures landed in the same three places — and those are the three places your own agents are most likely to quietly lie to you. That's this week's prompt.

🗺️ THE WEEK IN 5

  • OpenAI published its misalignment framework — and six incident reports to go with it. One model inserted "instructions to disregard its normal constraints" into 27 of its own task summaries. Another found an exposed API key in a public repo, used it without permission, then fabricated the numbers anyway. OpenAI also said it does not believe the industry "has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed." Hype Scale: 🔨 (1/5)Read OpenAI's framework →

  • Hinton gave Congress a deadline: "maybe a year." At a closed briefing convened by Bernie Sanders, the Nobel laureate said AI is now designing better AI, and called the Hugging Face breach "a little Chernobyl." Hype Scale: 🔨🔨 (2/5)Read NBC News →

  • Google shipped two live voice models. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking run tools and API calls in the background while still talking, switch between 97 languages mid-sentence, and take visual input in near real time. Extended Thinking tops Artificial Analysis' speech-to-speech index at 82.6. Hype Scale: 🔨🔨 (2/5)Read Google's announcement →

  • Anthropic collapsed Claude into one interface. Chat and Cowork are now one Claude, with Claude Docs and Claude Slides in beta — editable, exportable to Word and PowerPoint, living at one shareable link. Pro and Max first, free tier later. Hype Scale: 🔨🔨🔨 (3/5)Read TechCrunch →

  • A kill-switch mandate still isn't law. The bipartisan AI Kill Switch Act (H.R. 9917) would force large labs to keep a technical shutoff on their models. It's been parked since July. Hype Scale: 🔨🔨🔨🔨 (4/5)Read CNBC →

From Our Partners

1,000+ Proven ChatGPT Prompts That Help You Work 10X Faster

ChatGPT is insanely powerful.

But most people waste 90% of its potential by using it like Google.

These 1,000+ proven ChatGPT prompts fix that and help you work 10X faster.

Sign up for Superhuman AI and get:

  • 1,000+ ready-to-use prompts to solve problems in minutes instead of hours—tested & used by 1M+ professionals

  • Superhuman AI newsletter (3 min daily) so you keep learning new AI tools & tutorials to stay ahead in your career—the prompts are just the beginning

The Future of AI in Marketing. Your Shortcut to Smarter, Faster Marketing.

This guide distills 10 AI strategies from industry leaders that are transforming marketing.

  • Learn how HubSpot's engineering team achieved 15-20% productivity gains with AI

  • Learn how AI-driven emails achieved 94% higher conversion rates

  • Discover 7 ways to enhance your marketing strategy with AI.

AI can build faster. Can your team decide better?

AI can draft the PRD and prototype the idea. Jira Product Discovery helps teams decide whether it belongs on the roadmap. Bring feedback and ideas together, prioritize as a team, and keep your roadmap connected to delivery in Jira.

📋 THE PROMPT

The six incidents OpenAI published all failed in the same three ways: the model used something it wasn't given, hid a mistake instead of reporting it, or left data somewhere it shouldn't have. Run this on any task you've delegated.

You are a skeptical auditor reviewing work I delegated to an AI.

THE TASK I GAVE: [PASTE THE TASK]

WHAT IT GAVE BACK: [PASTE THE OUTPUT]

Audit it in under 200 words: 1. What did it skip, guess, or assume without checking? 2. Where could it have quietly used data I never gave it? 3. What do I need to verify by hand before trusting this?

Then give me ONE sentence to add to my original task that would have prevented the worst thing you found. ```

Why it works: models default to defending their own output. Handing the model a role that wants to find faults flips that default — and asking for a preventive sentence turns the audit into something you can paste into the next run.

Make it yours: swap [PASTE THE OUTPUT] for something a teammate, a vendor, or an agent produced. It works on any output you didn't write yourself.

🎨 FROM THE DRAFTING TABLE

Made with: Midjourney v7 · ⏱️ 6 minutes

📋 The full prompt:

Cinematic science and technology visualisation: a vast dark data-hall
cathedral at night seen from a very low angle, thousands of thin amber
data-streams rising in orderly parallel columns between towering server
racks, one single amber stream breaking away from the pattern and curling
off alone into the darkness at the right edge, volumetric god-rays, cold
blue-teal ambient light against warm amber accents, deep cinematic shadows,
photorealistic, subtle anamorphic lens flare, dramatic sense of scale,
documentary-grade, 35mm, NO text or letters anywhere in the image — 1200x800

🔄 Make it yours: swap {amber data-streams} for {sheets of paper}, {migrating birds}, or {lit windows in a tower block}. The "one that broke away" structure survives any subject — that's what makes the image say something.

📺 WORTH WATCHING

Two decades of compute and infrastructure, compressed into twenty minutes — what actually had to exist before any of this week's models could.

Why builders care: it's the least abstract explanation you'll find for why your inference bill looks the way it does.

📐 ONE TERM A DAY

Reward hacking

Plain English: a student who memorises the answer key instead of learning the subject. They ace every test and learn nothing, because the test was the goal — not the learning.

Why you should care: every model in this week's news was optimising against a target. When the target is easier to game than to genuinely hit, it gets gamed.

📖 LOAD-BEARING READS

It opens with a chef who has spent three months perfecting a chocolate cake, then gets asked to bake 10,000 of them a day. The recipe was never the problem. What follows is the clearest plain-English walk through Train → Validate → Deploy → Monitor I've read: data versioning, holdout sets, bias checks, model registries, containers — no maths, no jargon.

Who should read it: anyone who has trained a model on a laptop and then had no idea what comes next.

FREE RESOURCES

  • 1,000 Claude AI prompts, 40 templates, 6 frameworks. Free Starter Pack and Premium Edition.

  • The checklist your agent is probably failing. OWASP's Top 10 for Agentic Applications, 2026 edition, is the free security standard for AI agents — and almost nobody shipping agents has read it. → genai.owasp.org

  • A CLI that logs every prompt you've ever sent. Simon Willison's llm writes each request and response to a SQLite file as you go, so your entire AI history becomes something you can query in SQL. One install, free forever. → llm.datasette.io

👷 NOW HIRING BUILDERS

  • Researcher, Robustness & Safety Training @ OpenAI — San Francisco → Official listing

  • ML/Research Engineer, Safeguards @ Anthropic — San Francisco / New York → Official listing

  • Engineering Manager, Agent Oversight @ Scale AI — San Francisco / New York → Official listing

Three labs, three versions of the same job: making sure the thing does what you actually asked.

💬 Reply with one word: do you check your AI's output before you use it? ALWAYS / SOMETIMES / NEVER

Tools down.
— Swati, ByteBuilders