RT Ryan Greenblatt: Full automation of AI R&D likely yields a large speed up even if it doesn't get faster and faster (aka a software-only singularity...
RT Elizabeth Barnes: Re Sometimes people outside the field say things like “The AI situation can’t be that bad, there must be experts who are on top...
RT david rein: AI labs have started developing systems to monitor internally deployed AI agents for misaligned behavior. Earlier this year, I spent a ...
RT Ryan Greenblatt: Here are some of my top candidates for big pushes to do right now on technical AI safety (low effort notes). Much better model org...
RT METR: We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-hor...
It's often good that we have conservative epistemic norms in science. This makes it more challenging to address AI risk before it's too late, but not ...
RT Ben Snodin: 1/ New blog post where I try to figure out what makes tasks easy vs hard for AI agents, using @METR_Evals time horizon data. Short vers...
RT Helen Toner: Sometimes when people hear that the AI companies are trying to automate their own research, they think that just means using AI to wri...
RT Awni Hannun: Adopting Claude speak in my regular life, episode 1: Partner: Did you do the dishes tonight? Me: Yes they're done. Partner: Why are th...
RT David Shor: Re @jon_stokes If you control for "Knew what SlateStarCodex was in 2015" then financial exposure to AI labs is massively anti-correlate...
RT Kelsey Piper: I have a bunch of secret AI benchmarks I only reveal when they fall, and today one did. I give the AI 1000 words written by me and ne...
RT Yafah Edelman: We made a benchmark with METR that was incredibly hard and it is going to be saturated before we release it! Opus 4.6 probably has m...
RT david rein: Re In 2023 I made GPQA, and it saturated in about two years. Here, the benchmark I was working on probably saturated while we were maki...
RT METR: We co-developed MirrorCode with @EpochAIResearch to test AI on extremely long-horizon blackbox software reimplementation tasks. We found that...
RT Ryan Greenblatt: I tenatively believe it would be good if all AI companies had a policy of doing external deployment before internal deployment, be...
RT Joel Becker: in the alignment risk update for @AnthropicAI's mythos preview, ~all evaluations look good for anthropic except for @METR_Evals' best ...
RT Ryan Greenblatt: AIs are much better at easy-and-cheap-to-verify SWE tasks than I expected: I've seen AIs autonomously do perhaps 3-12 months of us...
RT Helen Toner: Everyone knows "AGI" doesn't have a single clear definition, but usually people try to fix that by proposing new ones. I wrote about w...
RT Ryan Greenblatt: This seems like a good breakdown, but: - I disagree quantitatively with the fake-data graphs: at parity I expect AIs to ~7x AI R&D...
RT Eli Lifland: AI timelines update: @DKokotajlo and I have updated our timelines earlier by ~1.5 years over the last 3 months, primarily due to (a) e...
RT Helen Toner: Andrej @karpathy on the inexorable logic of automating more and more of AI research: "To get the most of out the tools... you have to ...
RT Rohin Shah: "Just read the chain of thought" is one of our best safety techniques. Why does it work? Because models can only think opaquely for a s...