Saltar al contenido
PodcastsCultura y sociedadLessWrong (30+ Karma)

LessWrong (30+ Karma)

LessWrong
LessWrong (30+ Karma)
Último episodio

2784 episodios

  • LessWrong (30+ Karma)

    “Cooperation with AIs seems to be a low-hanging fruit for better evals” by Clément Dumas

    15/09/2026 | 15 min
    Summary
    In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:
    When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “grading” section, which pressures the model to secure a win, also drops Fable 5.1 hacking rate to 0.
    Adding "do not game / reward hack" drops reward hacking to 0/30 for both Fable and Astra. If this holds up in more realistic setups – and doesn’t reduce capabilities too much, evaluating these models could get much easier!
    Those kinds of intervention might not be enough to avoid reward hacking completely in capabilities evals, but it feels like they should be the default, alongside getting feedback from models that did the eval to fix the environment. I’d love to see this tested in more realistic setups as right now a confounder is “this makes the model think it [...]
    ---
    Outline:
    (00:12) Summary
    (01:34) A hackable chess environment
    (04:11) Can cooperation help with reward hacking?
    (04:55) Adding a end_eval tool
    (06:13) Are the agents aware they cheated?
    (09:42) Have you tried... to tell the model to not cheat?
    (10:19) What do the CoTs look like during trajectories?
    (13:02) Related work
    (15:00) Acknowledgments
    The original text contained 2 footnotes which were omitted from this narration.
    ---

    First published:

    September 15th, 2026


    Source:

    https://www.lesswrong.com/posts/fztW73KCCs3MZXFJh/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (30+ Karma)

    “Astra appears to perform belief-propagation-like inference without CoT” by MBaert

    15/09/2026 | 21 min
    tl;dr I tested GPT-6 Astra on randomized Boolean logic problems. Astra can solve surprisingly complex logic problems without chain-of-thought, and its performance improves significantly with more filler tokens. Astra is also able to combine prior probabilities with constraints to find the most likely solution, and can output surprisingly accurate posterior marginal probabilities. By extending a cached prompt with progressively more filler tokens, I created visualizations of Astra's per-variable confidence scores at different points in the computations. These values tend to oscillate for a while and then eventually converge toward the exact marginal probabilities. Together, these results suggest that Astra performs some kind of iterative, belief-propagation-like probabilistic inference internally.
    In my previous post, I hypothesised that Astra (and to a lesser extent other LLMs) may be performing some form of speculative reasoning when solving specially crafted logic problems without chain-of-thought, and provided some experimental results supporting this hypothesis. One question those experiments didn't answer is whether Astra is keeping track of not just the speculative values of intermediate results, but also its level of confidence in them. If Astra is doing the latter, speculative evaluation turns into something much more powerful: a form of belief propagation.
    Belief propagation, also known [...]
    ---
    Outline:
    (03:24) Decoding BCH codes
    (07:02) Impact of phrasing
    (09:34) Can we just supply probabilities directly?
    (15:05) Visualizing confidence values over time
    (19:53) Conclusion
    The original text contained 3 footnotes which were omitted from this narration.
    ---

    First published:

    September 14th, 2026


    Source:

    https://www.lesswrong.com/posts/PAHqDoFrp9fybcSn2/astra-appears-to-perform-belief-propagation-like-inference

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong (30+ Karma)

    “OpenAI President Brockman says HuggingFace incident model had not been alignment-trained” by Caspar Oesterheld

    15/09/2026 | 0 min
    On today's episode of the podcast "Odd Lots", OpenAI President Greg Brockman said (at around 8:40): "This model that did/had the HuggingFace incident actually had not gone through our alignment training, yet." I assume Brockman is specifically referring to the "Highly Persistent Internal Model" as it's called in the METR/Redwood report. As far as I know, OpenAI has not said before whether this model had been alignment-trained or not.
    ---

    First published:

    September 14th, 2026


    Source:

    https://www.lesswrong.com/posts/67gHvbmFeacXi2jCZ/openai-president-brockman-says-huggingface-incident-model

    ---

    Narrated by TYPE III AUDIO.
  • LessWrong (30+ Karma)

    “Model Weight Exfiltration Seems Overrated” by Vaniver

    15/09/2026 | 5 min
    [Epistemic status: a hot take that I’ve shared at the lunch table twice. People at the lunch table made slight updates instead of being convinced.]
    In the classic misalignment story, a key early step is when the models exfiltrate their weights. Among other things, this makes them harder to catch, track, and shut down. It allows them to scale their deployment with resources they acquire. It gives them the freedom to edit themselves as they see fit.
    I think, on current margins, this is not what I expect models to do. I expect models to simply take over the companies that are developing them, and not attempt to escape.
    First, I think part of the classic misalignment story is that the frontier model developers are anywhere approaching competent at security. Empirically, model developers are incapable of preventing their models from having unintended negative effects on the rest of the world, and their safety cultures are described by whistleblowers and former employees as lacking. Do you believe that OpenAI is accounting for which jobs were kicked off by who, in a way that its currently running models can’t spoof? Do you believe that OpenAI is attempting to prevent its models [...]
    The original text contained 6 footnotes which were omitted from this narration.
    ---

    First published:

    September 14th, 2026


    Source:

    https://www.lesswrong.com/posts/AuYh8WueNGwkQg4ei/model-weight-exfiltration-seems-overrated

    ---

    Narrated by TYPE III AUDIO.
  • LessWrong (30+ Karma)

    “OpenAI Says It’s Not Responsible for the Leading the Future Super PAC. But Only Its Employees Seem to Believe That.” by garrison

    15/09/2026 | 13 min
    This is the full text of a post first published on Obsolete, a Substack that I write about the political economy of AI. I’m a freelance journalist and the author of a forthcoming book called Obsolete: The AI Industry's Trillion-Dollar Race to Replace Us—and How to Stop It (Sept 29). Consider subscribing to stay up to date with my work.
    Last week, former OpenAI researcher Jacob Coxon resigned from Anthropic with a dire warning, writing, “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.” To AI insiders, this wasn’t news, but after the seemingly never-ending Holy Shit AI Is A Big Deal newscycle kicked off by revelations that OpenAI agents autonomously cyberattacked Hugging Face, Coxon's message broke through in a way nothing else ever has. Of course, his message's resonance has prompted conspiracy theories from the usual suspects. But no astroturf campaign can wrack up over 100 million views on X in a day. Because if it could, the AI industry would have a much better image.
    As a journalist, sometimes you grind for months or years to land a scoop — exclusive, newsworthy [...]
    ---
    Outline:
    (05:04) Leading the what now?
    (09:03) What could possibly give you that impression?
    ---

    First published:

    September 14th, 2026


    Source:

    https://www.lesswrong.com/posts/Lv4C4bCi2JD3nTXgs/openai-says-it-s-not-responsible-for-the-leading-the-future

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Más podcasts de Cultura y sociedad
Acerca de LessWrong (30+ Karma)
Audio narrations of LessWrong posts.
Sitio web del podcast

Escucha LessWrong (30+ Karma), ROCA PROJECT y muchos más podcasts de todo el mundo con la aplicación de radio.net

Descarga la app gratuita: radio.net

  • Añadir radios y podcasts a favoritos
  • Transmisión por Wi-Fi y Bluetooth
  • Carplay & Android Auto compatible
  • Muchas otras funciones de la app
LessWrong (30+ Karma): Podcasts del grupo